Wikitech
labswiki
https://wikitech.wikimedia.org/wiki/Main_Page
MediaWiki 1.47.0-wmf.20
first-letter
Media
Special
Talk
User
User talk
Wikitech
Wikitech talk
File
File talk
MediaWiki
MediaWiki talk
Template
Template talk
Help
Help talk
Category
Category talk
Obsolete
Obsolete talk
OfficeIT
OfficeIT talk
Tool
Tool talk
Nova Resource
Nova Resource Talk
Heira
Heira Talk
TimedText
TimedText talk
Module
Module talk
Deployments
0
4108
2458700
2458686
2026-09-19T16:22:52Z
ScheduleDeploymentBot
37566
Add [[gerrit:1343100]] to Monday, September 21 UTC morning backport window
2458700
wikitext
text/x-wiki
{{Navigation MediaWiki deployment}}
This page tracks '''upcoming''' '''deployments''' of software to the [[:m:Special:SiteMatrix|Wikimedia Foundation servers]].
== Getting started ==
Ensure you joined the {{irc|wikimedia-operations}} IRC channel as all deployment-related communications happen there.
If you need help, contact [[:mw:Wikimedia Release Engineering Team|Release Engineering]] on IRC at {{irc|wikimedia-releng}}; and ping Tyler (<code>thcipriani</code>).
* '''MediaWiki is deployed weekly''' through the [[/Train|Deployment Train]]. Other services follow their own schedule.
* '''Times are pinned to San Francisco''', thus the UTC time changes in March and November per [[:en:Daylight saving time in the United States|DST]].
* '''Prefer regular [[Backport windows]]''' over adding new windows. To request deployment of a config change or backport, add your username and Gerrit URL to one of the backport windows on this page. You must be online in #wikimedia-operations on IRC during your deployment and install [[WikimediaDebug]] ahead of time. The #wikimedia-operations channel requires you to [[:m:IRC/Instructions#Register your nickname, identify, and enforce|register your nickname]] before you can join.
** You can use the '''backport scheduling tool''' to more easily edit this page: <div style="text-align: center; margin: 1em 0">{{Clickable button 2|:toollabs:schedule-deployment|Schedule a backport|class=mw-ui-progressive}}</div>
* Tasks that meet [[/Inclusion criteria|Inclusion criteria]] '''require their own windows''', which includes long-running tasks. '''Schedule more time''' than you think you need to account for delays and set backs, we recommend one hour for most tasks.
**To create or modify a recurring deploy window, send a patchset to [[:gitlab:repos/releng/release/-/blob/main/make-deployment-calendar/deployments-calendar.yaml|deployments-calendar.yaml file]] in <code>repos/releng/release.git</code>.
**To create an one-off window, simply edit this page accordingly
** '''Announce''' changes to the [[mail:ops|ops mailing list]] ahead of time if you anticipate or are uncertain about noticeable impacts to database load, HTTP caching, or the introduction of new cookies.
** '''Announce''' deployments of major features to the community via [[:m:Tech/News/Next|Tech News]] and/or via other [[:mw:Wikimedia_Product_Guidance/Communication_channels|Product communication channels]].
* '''Something went wrong?''' See [[Incident response]]. Is there a user-impacting problem? Communicate in the {{irc|wikimedia-operations}} IRC channel. If there is a Phabricator task, ensure [[:phab:tag/wikimedia-incident/|#Wikimedia-Incident]] is tagged, and consider setting the [[:mw:Phabricator/Project_management#Priority_levels|Unbreak Now]] priority.
__TOC__
{{anchor|Next Week|Near Term|Near term|Near-term}}{{clear}}
[[Category:Deployment]]
{{Note|content=Subscribe in Google Calendar via <code>wikimedia.org_rudis09ii2mm5fk4hgdjeh1u64@group.calendar.google.com</code>.<br>This may not include one-off windows. '''If there are differences, then the wiki page is canonical and correct'''.}}
== Week of September 21 ==
=== {{Deployment_day|date=2026-09-23}} ===
{{Deployment calendar event card
|when=2026-09-23 13:00 UTC
|length=4
|window=Datacenter switch over
|who={{ircnick|slyngs}}
|what=Datacenter switch over
}}
==Week of September 21==
==={{Deployment_day|date=2026-09-20}}===
{{Deployment calendar event card
|when=2026-09-20 00:00 SF
|length=24
|window=No deploys all day! See [[Deployments/Emergencies]] if things are broken.
|who=
|what=No Deploys
}}
==={{Deployment_day|date=2026-09-21}}===
{{Deployment calendar event card
|when=2026-09-21 00:00 SF
|length=1
|window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}}
|what={{ircnick|Ameisenigel|Ameisenigel}}
{{deploy|type=config|gerrit=1343100|title=Disable wgMFCustomSiteModules on German Wikipedia|status=}} - {{phabricator|T403380}}
{{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-21 02:00 UTC
|length=1
|window=Automatic deployment of MediaWiki to pretrain wikis - see [[mw:Pretrain]]
|who=N/A
|what=Deploy <code>wmf/next</code> to pretrain environment
}}
{{Deployment calendar event card
|when=2026-09-21 03:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-21 06:00 SF
|length=1
|window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}
|what={{ircnick|codenamenoreste|Codename Noreste}}
{{deploy|type=config|gerrit=1342825|title=eswiki: Add abusefilter-access-protected-vars to abusefilter user group|status=}} - {{phabricator|T436652}}
{{ircnick|mfossati|Marco}}
{{deploy|type=1.47.0-wmf.20|gerrit=1343122|title=Let AA measure eligible readers w/o beta opt-in|status=}} - {{phabricator|T437076}}
{{ircnick|Dreamy_Jazz|WBrown (WMF)}}
{{deploy|type=config|gerrit=1337580|title=nlwiki: enable SecurePoll local elections|status=}} - {{phabricator|T434045}}
{{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-21 07:30 SF
|length=0.5
|window=Test Kitchen Experiment Deployment Window
|who=Test Kitchen
|what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]].
}}
{{Deployment calendar event card
|when=2026-09-21 08:30 SF
|length=0.5
|window=Wikimedia Portals Update
|who={{ircnick|jan_drewniak|Jan Drewniak}}
|what=Weekly window for the portals page: https://www.wikipedia.org/
}}
{{Deployment calendar event card
|when=2026-09-21 10:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-21 10:00 SF
|length=0.5
|window=Wikidata Query Service weekly deploy
|who={{ircnick|ryankemper|Ryan}}
|what=...
}}
{{Deployment calendar event card
|when=2026-09-21 13:00 SF
|length=1
|window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-21 14:00 SF
|length=2
|window=Weekly Security deployment window
|who={{ircnick|alexsanford|Alex}}, {{ircnick|Reedy|Sam}}, {{ircnick|sbassett|Scott}}, {{ircnick|Maryum|Maryum}}, {{ircnick|manfredi|Manfredi}}
|what=Held deployment window for Security-team related deploys.
}}
{{Deployment calendar event card
|when=2026-09-21 16:00 SF
|length=1
|window=Readers deployment window
|who=Readers
|what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start
}}
{{Deployment calendar event card
|when=2026-09-21 19:00 SF
|length=1
|window=Automatic branching of MediaWiki, extensions, skins, and vendor – see [[Heterogeneous deployment/Train deploys]]
|who=N/A
|what=Branch <code>wmf/0.00.0-wmf.0</code>
}}
{{Deployment calendar event card
|when=2026-09-21 20:00 SF
|length=1
|window=Automatic deployment of MediaWiki, extensions, skins, and vendor to testwikis only – see [[Heterogeneous deployment/Train deploys]]
|who=N/A
|what=Deploy <code>wmf/0.00.0-wmf.0</code> to testwikis
}}
{{Deployment calendar event card
|when=2026-09-21 21:00 SF
|length=1
|window=Automatic removal of all obsolete MediaWiki versions from the deployment and bare metal servers (except the most-recent obsolete version)
|who=N/A
|what=Runs <code>scap clean auto</code>
}}
{{Deployment calendar event card
|when=2026-09-21 23:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-21 23:00 SF
|length=0.5
|window=Primary database switchover
|who={{ircnick|marostegui|Manuel Arostegui}}, {{ircnick|cezmunsta|Ceri Williams}}, {{ircnick|federico3|Federico Ceratto}}
|what=Held deployment window for database primary masters maintenance
}}
==={{Deployment_day|date=2026-09-22}}===
{{Deployment calendar event card
|when=2026-09-22 00:00 SF
|length=1
|window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-22 02:00 UTC
|length=1
|window=Automatic deployment of MediaWiki to pretrain wikis - see [[mw:Pretrain]]
|who=N/A
|what=Deploy <code>wmf/next</code> to pretrain environment
}}
{{Deployment calendar event card
|when=2026-09-22 03:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-22 05:00 SF
|length=1
|window=Mobileapps/RESTBase/Wikifeeds
|who=Content Transform Team
|what=Content transform team node services (mobileapps/wikifeeds)
}}
{{Deployment calendar event card
|when=2026-09-22 06:00 SF
|length=1
|window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-22 07:00 SF
|length=0.5
|window=Test Kitchen UI Deployment Window
|who=Experimentation Platform Team
|what=Deployment of Test Kitchen UI (fka MPIC)
}}
{{Deployment calendar event card
|when=2026-09-22 07:30 SF
|length=0.5
|window=Test Kitchen Experiment Deployment Window
|who=Test Kitchen
|what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]].
}}
{{Deployment calendar event card
|when=2026-09-22 08:00 SF
|length=1
|window=SRE Collaboration Services office hours
|who={{ircnick|jelto|Jelto}}, {{ircnick|arnoldokoth|Arnold}}, {{ircnick|mutante|Daniel}}, {{ircnick|arnaudb|Arnaud}}
|what=Services including Gerrit, Phorge (Phabricator), GitLab
}}
{{Deployment calendar event card
|when=2026-09-22 09:00 SF
|length=1
|window=[[Puppet request window]]<br/><small>'''(Max 6 patches)'''</small>
|who={{ircnick|jhathaway|JHathaway}}, {{ircnick|rzl|Reuven}}
|what={{ircnick|zabe|Zabe}}
* {{gerrit|1339694}} Add Apache configuration for wikipedia-ar-arbcom.wikimedia.org
{{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to Puppet change''
}}
{{Deployment calendar event card
|when=2026-09-22 10:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-22 11:00 SF
|length=2
|window=MediaWiki train - Utc-7 Version
|who={{ircnick|thcipriani|Tyler}}, {{ircnick|thcipriani|Tyler}}
|what=[[mw:MediaWiki 1.00/Roadmap#Schedule for the deployments|1.00 schedule]]
{{DeployOneWeekMini|0.00.0-wmf.0->0.00.0-wmf.0|0.00.0-wmf.0|0.00.0-wmf.0}}
* group0 to [[mw:MediaWiki_1.00/wmf.0|0.00.0-wmf.0]]
* '''Blockers: {{phabricator|None}}'''
}}
{{Deployment calendar event card
|when=2026-09-22 13:00 SF
|length=1
|window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-22 14:00 SF
|length=1
|window=Readers deployment window
|who=Readers
|what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start
}}
{{Deployment calendar event card
|when=2026-09-22 23:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
==={{Deployment_day|date=2026-09-23}}===
{{Deployment calendar event card
|when=2026-09-23 00:00 SF
|length=1
|window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-23 02:00 UTC
|length=1
|window=Automatic deployment of MediaWiki to pretrain wikis - see [[mw:Pretrain]]
|who=N/A
|what=Deploy <code>wmf/next</code> to pretrain environment
}}
{{Deployment calendar event card
|when=2026-09-23 03:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-23 04:00 SF
|length=1
|window=[[mw:Services|Services]] – [[Citoid]] / [[Zotero]]
|who=Marielle ({{ircnick|mvolz}})
|what=See [[mw:Citoid|Citoid]]
}}
{{Deployment calendar event card
|when=2026-09-23 06:00 SF
|length=1
|window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-23 07:00 SF
|length=1
|window=Wikifunctions Services UTC Afternoon
|who=Abstract Wikipedia team (Africa, Europe, Eastern Americas)
|what=Wikifunctions back-end k8s services
}}
{{Deployment calendar event card
|when=2026-09-23 07:30 SF
|length=0.5
|window=Test Kitchen Experiment Deployment Window
|who=Test Kitchen
|what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]].
}}
{{Deployment calendar event card
|when=2026-09-23 10:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-23 11:00 SF
|length=2
|window=MediaWiki train - Utc-7 Version
|who={{ircnick|thcipriani|Tyler}}, {{ircnick|thcipriani|Tyler}}
|what=[[mw:MediaWiki 1.00/Roadmap#Schedule for the deployments|1.00 schedule]]
{{DeployOneWeekMini|0.00.0-wmf.0|0.00.0-wmf.0->0.00.0-wmf.0|0.00.0-wmf.0}}
* group1 to [[mw:MediaWiki_1.00/wmf.0|0.00.0-wmf.0]]
* '''Blockers: {{phabricator|None}}'''
}}
{{Deployment calendar event card
|when=2026-09-23 13:00 SF
|length=1
|window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-23 14:00 SF
|length=1
|window=Wikifunctions Services UTC Late
|who=Abstract Wikipedia team (North and South America)
|what=Wikifunctions back-end k8s services
}}
{{Deployment calendar event card
|when=2026-09-23 15:00 SF
|length=1
|window=Readers deployment window
|who=Readers
|what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start
}}
{{Deployment calendar event card
|when=2026-09-23 23:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-23 23:00 SF
|length=0.5
|window=Primary database switchover
|who={{ircnick|marostegui|Manuel Arostegui}}, {{ircnick|cezmunsta|Ceri Williams}}, {{ircnick|federico3|Federico Ceratto}}
|what=Held deployment window for database primary masters maintenance
}}
==={{Deployment_day|date=2026-09-24}}===
{{Deployment calendar event card
|when=2026-09-24 00:00 SF
|length=1
|window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-24 02:00 UTC
|length=1
|window=Automatic deployment of MediaWiki to pretrain wikis - see [[mw:Pretrain]]
|who=N/A
|what=Deploy <code>wmf/next</code> to pretrain environment
}}
{{Deployment calendar event card
|when=2026-09-24 03:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-24 05:00 SF
|length=1
|window=Mobileapps/RESTBase/Wikifeeds
|who=Content Transform Team
|what=Content transform team node services (mobileapps/wikifeeds)
}}
{{Deployment calendar event card
|when=2026-09-24 06:00 SF
|length=1
|window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-24 07:30 SF
|length=0.5
|window=Test Kitchen Experiment Deployment Window
|who=Test Kitchen
|what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]].
}}
{{Deployment calendar event card
|when=2026-09-24 09:00 SF
|length=1
|window=[[Puppet request window]]<br/><small>'''(Max 6 patches)'''</small>
|who={{ircnick|jhathaway|JHathaway}}, {{ircnick|rzl|Reuven}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to Puppet change''
}}
{{Deployment calendar event card
|when=2026-09-24 10:00 SF
|length=1
|window=Cloud Services/Technical Documentation weekly deploy (Toolhub, Developer portal, Striker)
|who={{ircnick|bd808}}
|what=...
}}
{{Deployment calendar event card
|when=2026-09-24 10:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-24 11:00 SF
|length=2
|window=MediaWiki train - Utc-7 Version
|who={{ircnick|thcipriani|Tyler}}, {{ircnick|thcipriani|Tyler}}
|what=[[mw:MediaWiki 1.00/Roadmap#Schedule for the deployments|1.00 schedule]]
{{DeployOneWeekMini|0.00.0-wmf.0|0.00.0-wmf.0|0.00.0-wmf.0->0.00.0-wmf.0}}
* group2 to [[mw:MediaWiki_1.00/wmf.0|0.00.0-wmf.0]]
* '''Blockers: {{phabricator|None}}'''
}}
{{Deployment calendar event card
|when=2026-09-24 13:00 SF
|length=1
|window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-24 14:00 SF
|length=1
|window=Readers deployment window
|who=Readers
|what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start
}}
{{Deployment calendar event card
|when=2026-09-24 23:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
==={{Deployment_day|date=2026-09-25}}===
{{Deployment calendar event card
|when=2026-09-25 00:00 SF
|length=24
|window=No deploys all day! See [[Deployments/Emergencies]] if things are broken.
|who=
|what=No Deploys
}}
{{Deployment calendar event card
|when=2026-09-25 04:00 SF
|length=0.5
|window=GitLab version upgrades
|who={{ircnick|jelto|Jelto}}, {{ircnick|arnoldokoth|Arnold}}, {{ircnick|mutante|Daniel}}, {{ircnick|arnaudb|Arnaud}}
|what=GitLab version upgrades
}}
==={{Deployment_day|date=2026-09-26}}===
{{Deployment calendar event card
|when=2026-09-26 00:00 SF
|length=24
|window=No deploys all day! See [[Deployments/Emergencies]] if things are broken.
|who=
|what=No Deploys
}}
==Week of September 28==
==={{Deployment_day|date=2026-09-27}}===
{{Deployment calendar event card
|when=2026-09-27 00:00 SF
|length=24
|window=No deploys all day! See [[Deployments/Emergencies]] if things are broken.
|who=
|what=No Deploys
}}
==={{Deployment_day|date=2026-09-28}}===
{{Deployment calendar event card
|when=2026-09-28 00:00 SF
|length=1
|window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-28 02:00 UTC
|length=1
|window=Automatic deployment of MediaWiki to pretrain wikis - see [[mw:Pretrain]]
|who=N/A
|what=Deploy <code>wmf/next</code> to pretrain environment
}}
{{Deployment calendar event card
|when=2026-09-28 03:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-28 06:00 SF
|length=1
|window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-28 07:30 SF
|length=0.5
|window=Test Kitchen Experiment Deployment Window
|who=Test Kitchen
|what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]].
}}
{{Deployment calendar event card
|when=2026-09-28 08:30 SF
|length=0.5
|window=Wikimedia Portals Update
|who={{ircnick|jan_drewniak|Jan Drewniak}}
|what=Weekly window for the portals page: https://www.wikipedia.org/
}}
{{Deployment calendar event card
|when=2026-09-28 10:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-28 10:00 SF
|length=0.5
|window=Wikidata Query Service weekly deploy
|who={{ircnick|ryankemper|Ryan}}
|what=...
}}
{{Deployment calendar event card
|when=2026-09-28 13:00 SF
|length=1
|window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-28 14:00 SF
|length=2
|window=Weekly Security deployment window
|who={{ircnick|alexsanford|Alex}}, {{ircnick|Reedy|Sam}}, {{ircnick|sbassett|Scott}}, {{ircnick|Maryum|Maryum}}, {{ircnick|manfredi|Manfredi}}
|what=Held deployment window for Security-team related deploys.
}}
{{Deployment calendar event card
|when=2026-09-28 16:00 SF
|length=1
|window=Readers deployment window
|who=Readers
|what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start
}}
{{Deployment calendar event card
|when=2026-09-28 19:00 SF
|length=1
|window=Automatic branching of MediaWiki, extensions, skins, and vendor – see [[Heterogeneous deployment/Train deploys]]
|who=N/A
|what=Branch <code>wmf/1.47.0-wmf.22</code>
}}
{{Deployment calendar event card
|when=2026-09-28 20:00 SF
|length=1
|window=Automatic deployment of MediaWiki, extensions, skins, and vendor to testwikis only – see [[Heterogeneous deployment/Train deploys]]
|who=N/A
|what=Deploy <code>wmf/1.47.0-wmf.22</code> to testwikis
}}
{{Deployment calendar event card
|when=2026-09-28 21:00 SF
|length=1
|window=Automatic removal of all obsolete MediaWiki versions from the deployment and bare metal servers (except the most-recent obsolete version)
|who=N/A
|what=Runs <code>scap clean auto</code>
}}
{{Deployment calendar event card
|when=2026-09-28 23:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-28 23:00 SF
|length=0.5
|window=Primary database switchover
|who={{ircnick|marostegui|Manuel Arostegui}}, {{ircnick|cezmunsta|Ceri Williams}}, {{ircnick|federico3|Federico Ceratto}}
|what=Held deployment window for database primary masters maintenance
}}
==={{Deployment_day|date=2026-09-29}}===
{{Deployment calendar event card
|when=2026-09-29 00:00 SF
|length=1
|window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-29 02:00 UTC
|length=1
|window=Automatic deployment of MediaWiki to pretrain wikis - see [[mw:Pretrain]]
|who=N/A
|what=Deploy <code>wmf/next</code> to pretrain environment
}}
{{Deployment calendar event card
|when=2026-09-29 03:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-29 05:00 SF
|length=1
|window=Mobileapps/RESTBase/Wikifeeds
|who=Content Transform Team
|what=Content transform team node services (mobileapps/wikifeeds)
}}
{{Deployment calendar event card
|when=2026-09-29 06:00 SF
|length=1
|window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-29 07:00 SF
|length=0.5
|window=Test Kitchen UI Deployment Window
|who=Experimentation Platform Team
|what=Deployment of Test Kitchen UI (fka MPIC)
}}
{{Deployment calendar event card
|when=2026-09-29 07:30 SF
|length=0.5
|window=Test Kitchen Experiment Deployment Window
|who=Test Kitchen
|what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]].
}}
{{Deployment calendar event card
|when=2026-09-29 08:00 SF
|length=1
|window=SRE Collaboration Services office hours
|who={{ircnick|jelto|Jelto}}, {{ircnick|arnoldokoth|Arnold}}, {{ircnick|mutante|Daniel}}, {{ircnick|arnaudb|Arnaud}}
|what=Services including Gerrit, Phorge (Phabricator), GitLab
}}
{{Deployment calendar event card
|when=2026-09-29 09:00 SF
|length=1
|window=[[Puppet request window]]<br/><small>'''(Max 6 patches)'''</small>
|who={{ircnick|jhathaway|JHathaway}}, {{ircnick|rzl|Reuven}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to Puppet change''
}}
{{Deployment calendar event card
|when=2026-09-29 10:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-29 11:00 SF
|length=2
|window=MediaWiki train - Utc-7 Version
|who={{ircnick|brennen|Brennen}}, {{ircnick|jeena|Jeena}}
|what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]]
{{DeployOneWeekMini|1.47.0-wmf.21->1.47.0-wmf.22|1.47.0-wmf.21|1.47.0-wmf.21}}
* group0 to [[mw:MediaWiki_1.47/wmf.22|1.47.0-wmf.22]]
* '''Blockers: {{phabricator|T438218}}'''
}}
{{Deployment calendar event card
|when=2026-09-29 13:00 SF
|length=1
|window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-29 14:00 SF
|length=1
|window=Readers deployment window
|who=Readers
|what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start
}}
{{Deployment calendar event card
|when=2026-09-29 23:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
==={{Deployment_day|date=2026-09-30}}===
{{Deployment calendar event card
|when=2026-09-30 00:00 SF
|length=1
|window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-30 02:00 UTC
|length=1
|window=Automatic deployment of MediaWiki to pretrain wikis - see [[mw:Pretrain]]
|who=N/A
|what=Deploy <code>wmf/next</code> to pretrain environment
}}
{{Deployment calendar event card
|when=2026-09-30 03:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-30 04:00 SF
|length=1
|window=[[mw:Services|Services]] – [[Citoid]] / [[Zotero]]
|who=Marielle ({{ircnick|mvolz}})
|what=See [[mw:Citoid|Citoid]]
}}
{{Deployment calendar event card
|when=2026-09-30 06:00 SF
|length=1
|window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-30 07:00 SF
|length=1
|window=Wikifunctions Services UTC Afternoon
|who=Abstract Wikipedia team (Africa, Europe, Eastern Americas)
|what=Wikifunctions back-end k8s services
}}
{{Deployment calendar event card
|when=2026-09-30 07:30 SF
|length=0.5
|window=Test Kitchen Experiment Deployment Window
|who=Test Kitchen
|what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]].
}}
{{Deployment calendar event card
|when=2026-09-30 10:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-30 11:00 SF
|length=2
|window=MediaWiki train - Utc-7 Version
|who={{ircnick|brennen|Brennen}}, {{ircnick|jeena|Jeena}}
|what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]]
{{DeployOneWeekMini|1.47.0-wmf.22|1.47.0-wmf.21->1.47.0-wmf.22|1.47.0-wmf.21}}
* group1 to [[mw:MediaWiki_1.47/wmf.22|1.47.0-wmf.22]]
* '''Blockers: {{phabricator|T438218}}'''
}}
{{Deployment calendar event card
|when=2026-09-30 13:00 SF
|length=1
|window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-30 14:00 SF
|length=1
|window=Wikifunctions Services UTC Late
|who=Abstract Wikipedia team (North and South America)
|what=Wikifunctions back-end k8s services
}}
{{Deployment calendar event card
|when=2026-09-30 15:00 SF
|length=1
|window=Readers deployment window
|who=Readers
|what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start
}}
{{Deployment calendar event card
|when=2026-09-30 23:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-30 23:00 SF
|length=0.5
|window=Primary database switchover
|who={{ircnick|marostegui|Manuel Arostegui}}, {{ircnick|cezmunsta|Ceri Williams}}, {{ircnick|federico3|Federico Ceratto}}
|what=Held deployment window for database primary masters maintenance
}}
==={{Deployment_day|date=2026-10-01}}===
{{Deployment calendar event card
|when=2026-10-01 00:00 SF
|length=1
|window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-10-01 02:00 UTC
|length=1
|window=Automatic deployment of MediaWiki to pretrain wikis - see [[mw:Pretrain]]
|who=N/A
|what=Deploy <code>wmf/next</code> to pretrain environment
}}
{{Deployment calendar event card
|when=2026-10-01 03:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-10-01 05:00 SF
|length=1
|window=Mobileapps/RESTBase/Wikifeeds
|who=Content Transform Team
|what=Content transform team node services (mobileapps/wikifeeds)
}}
{{Deployment calendar event card
|when=2026-10-01 06:00 SF
|length=1
|window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-10-01 07:30 SF
|length=0.5
|window=Test Kitchen Experiment Deployment Window
|who=Test Kitchen
|what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]].
}}
{{Deployment calendar event card
|when=2026-10-01 09:00 SF
|length=1
|window=[[Puppet request window]]<br/><small>'''(Max 6 patches)'''</small>
|who={{ircnick|jhathaway|JHathaway}}, {{ircnick|rzl|Reuven}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to Puppet change''
}}
{{Deployment calendar event card
|when=2026-10-01 10:00 SF
|length=1
|window=Cloud Services/Technical Documentation weekly deploy (Toolhub, Developer portal, Striker)
|who={{ircnick|bd808}}
|what=...
}}
{{Deployment calendar event card
|when=2026-10-01 10:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-10-01 11:00 SF
|length=2
|window=MediaWiki train - Utc-7 Version
|who={{ircnick|brennen|Brennen}}, {{ircnick|jeena|Jeena}}
|what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]]
{{DeployOneWeekMini|1.47.0-wmf.22|1.47.0-wmf.22|1.47.0-wmf.21->1.47.0-wmf.22}}
* group2 to [[mw:MediaWiki_1.47/wmf.22|1.47.0-wmf.22]]
* '''Blockers: {{phabricator|T438218}}'''
}}
{{Deployment calendar event card
|when=2026-10-01 13:00 SF
|length=1
|window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-10-01 14:00 SF
|length=1
|window=Readers deployment window
|who=Readers
|what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start
}}
{{Deployment calendar event card
|when=2026-10-01 23:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
==={{Deployment_day|date=2026-10-02}}===
{{Deployment calendar event card
|when=2026-10-02 00:00 SF
|length=24
|window=No deploys all day! See [[Deployments/Emergencies]] if things are broken.
|who=
|what=No Deploys
}}
{{Deployment calendar event card
|when=2026-10-02 04:00 SF
|length=0.5
|window=GitLab version upgrades
|who={{ircnick|jelto|Jelto}}, {{ircnick|arnoldokoth|Arnold}}, {{ircnick|mutante|Daniel}}, {{ircnick|arnaudb|Arnaud}}
|what=GitLab version upgrades
}}
==={{Deployment_day|date=2026-10-03}}===
{{Deployment calendar event card
|when=2026-10-03 00:00 SF
|length=24
|window=No deploys all day! See [[Deployments/Emergencies]] if things are broken.
|who=
|what=No Deploys
}}
5jpsd6qckkomsuqlbnnpbytzujrkutk
Server Admin Log/Archives
0
4673
2458707
2455997
2026-09-19T20:09:01Z
JrandWP
37706
/* 2025-present */ archive
2458707
wikitext
text/x-wiki
<noinclude>{{process header
|previous=← [[Server Admin Log]]
|title=Server Admin Log
|section=(archives)
}}</noinclude>
<inputbox>
type=fulltext
prefix=Server Admin Log/
searchbuttonlabel=Search archives
break=no
</inputbox><noinclude>
==Archives==
</noinclude>
===2000s===
<div style="column-count:2;-moz-column-count:2;-webkit-column-count:2">
* [[Server Admin Log/Archive 1|Archive 1: 2004 Jun - 2004 Sep]]
* [[Server Admin Log/Archive 2|Archive 2: 2004 Oct - 2004 Nov]]
* [[Server Admin Log/Archive 3|Archive 3: 2004 Dec - 2005 Mar]]
* [[Server Admin Log/Archive 4|Archive 4: 2005 Apr - 2005 Jul]]
* [[Server Admin Log/Archive 5|Archive 5: 2005 Aug - 2005 Oct]], <small>with revision history 2004-06-23 to 2005-11-25</small>
* [[Server Admin Log/Archive 6|Archive 6: 2005 Nov - 2006 Feb]]
* [[Server Admin Log/Archive 7|Archive 7: 2006 Mar - 2006 Jun]]
* [[Server Admin Log/Archive 8|Archive 8: 2006 Jul - 2006 Sep]]
* [[Server Admin Log/Archive 9|Archive 9: 2006 Oct - 2007 Jan]], <small>with revision history 2005-11-25 to 2007-02-21</small>
* [[Server Admin Log/Archive 10|Archive 10: 2007 Feb - 2007 Jun]]
* [[Server Admin Log/Archive 11|Archive 11: 2007 Jul - 2007 Dec]]
* [[Server Admin Log/Archive 12|Archive 12: 2008 Jan - 2008 Jul]]
* [[Server Admin Log/2008-08|Archive 12a: 2008 Aug]]
* [[Server Admin Log/2008-09|Archive 12b: 2008 Sept]]
* [[Server Admin Log/Archive 13|Archive 13: 2008 Oct - 2009 Jun]]
* [[Server Admin Log/Archive 14|Archive 14: 2009 Jun - 2009 Dec]]
</div>
===2010s===
<div style="column-count:2;-moz-column-count:2;-webkit-column-count:2">
* [[Server Admin Log/Archive 15|Archive 15: 2010 Jan - 2010 Jun]]
* [[Server Admin Log/Archive 16|Archive 16: 2010 Jul - 2010 Oct]]
* [[Server Admin Log/Archive 17|Archive 17: 2010 Nov - 2010 Dec]]
* [[Server Admin Log/Archive 18|Archive 18: 2011 Jan - 2011 Jun]]
* [[Server Admin Log/Archive 19|Archive 19: 2011 Jul - 2011 Dec]]
* [[Server Admin Log/Archive 20|Archive 20: 2011 Dec - 2012 Jun]], <small>with revision history 2007-02-21 to 2012-03-27</small>
* [[Server Admin Log/Archive 21|Archive 21: 2012 Jul - 2013 Jan]]
* [[Server Admin Log/Archive 22|Archive 22: 2013 Jan - 2013 Jul]]
* [[Server Admin Log/Archive 23|Archive 23: 2013 Aug - 2013 Dec]]
* [[Server Admin Log/Archive 24|Archive 24: 2014 Jan - 2014 Mar]]
* [[Server Admin Log/Archive 25|Archive 25: 2014 April - 2014 September]]
* [[Server Admin Log/Archive 26|Archive 26: 2014 October - 2014 December]]
* [[Server Admin Log/Archive 27|Archive 27: 2015 January - 2015 July]]
* [[Server Admin Log/Archive 28|Archive 28: 2015 August - 2015 December]]
* [[Server Admin Log/Archive 29|Archive 29: 2016 January - 2016 May]]
* [[Server Admin Log/Archive 30|Archive 30: 2016 June - 2016 August]]
* [[Server Admin Log/Archive 31|Archive 31: 2016 September - 2016 December]]
* [[Server Admin Log/Archive 32|Archive 32: 2017 January - 2017 July]]
* [[Server Admin Log/Archive 33|Archive 33: 2017 August - 2017 December]]
* [[Server Admin Log/Archive 34|Archive 34: 2018 January - 2018 April]]
* [[Server Admin Log/Archive 35|Archive 35: 2018 May - 2018 August]]
* [[Server Admin Log/Archive 36|Archive 36: 2018 September - 2018 December]]
* [[Server Admin Log/Archive 37|Archive 37: 2019 January - 2019 April]]
* [[Server Admin Log/Archive 38|Archive 38: 2019 May - 2019 August]]
* [[Server Admin Log/Archive 39|Archive 39: 2019 September - 2019 December]]
</div>
===2020-2024===
<div style="column-count:2;-moz-column-count:2;-webkit-column-count:2">
* [[Server Admin Log/Archive 40|Archive 40: 2020 January - 2020 April]]
* [[Server Admin Log/Archive 41|Archive 41: 2020 May - 2020 July]]
* [[Server Admin Log/Archive 42|Archive 42: 2020 August - 2020 November]]
* [[Server Admin Log/Archive 43|Archive 43: 2020 December]]
* [[Server Admin Log/Archive 44|Archive 44: 2021 January - 2021 April]]
* [[Server Admin Log/Archive 45|Archive 45: 2021 May - 2021 July]]
* [[Server Admin Log/Archive 46|Archive 46: 2021 August - 2021 October]]
* [[Server Admin Log/Archive 47|Archive 47: 2021 November - 2021 December]]
* [[Server Admin Log/Archive 48|Archive 48: 2022 January]]
* [[Server Admin Log/Archive 49|Archive 49: 2022 February]]
* [[Server Admin Log/Archive 50|Archive 50: 2022 March]]
* [[Server Admin Log/Archive 51|Archive 51: 2022 April 1-15]]
* [[Server Admin Log/Archive 52|Archive 52: 2022 April 16-30]]
* [[Server Admin Log/Archive 53|Archive 53: 2022 May]]
* [[Server Admin Log/Archive 54|Archive 54: 2022 June]]
* [[Server Admin Log/Archive 55|Archive 55: 2022 July]]
* [[Server Admin Log/Archive 56|Archive 56: 2022 August]]
* [[Server Admin Log/Archive 57|Archive 57: 2022 September]]
* [[Server Admin Log/Archive 58|Archive 58: 2022 October]]
* [[Server Admin Log/Archive 59|Archive 59: 2022 November 1-15]]
* [[Server Admin Log/Archive 60|Archive 60: 2022 November 16-30]]
* [[Server Admin Log/Archive 61|Archive 61: 2022 December]]
* [[Server Admin Log/Archive 62|Archive 62: 2023 January]]
* [[Server Admin Log/Archive 63|Archive 63: 2023 February]]
* [[Server Admin Log/Archive 64|Archive 64: 2023 March]]
* [[Server Admin Log/Archive 65|Archive 65: 2023 April]]
* [[Server Admin Log/Archive 66|Archive 66: 2023 May]]
* [[Server Admin Log/Archive 67|Archive 67: 2023 June]]
* [[Server Admin Log/Archive 68|Archive 68: 2023 July]]
* [[Server Admin Log/Archive 69|Archive 69: 2023 August 1-15]]
* [[Server Admin Log/Archive 70|Archive 70: 2023 August 16-31]]
* [[Server Admin Log/Archive 71|Archive 71: 2023 September]]
* [[Server Admin Log/Archive 72|Archive 72: 2023 October]]
* [[Server Admin Log/Archive 73|Archive 73: 2023 November]]
* [[Server Admin Log/Archive 74|Archive 74: 2023 December]]
* [[Server Admin Log/Archive 75|Archive 75: 2024 January]]
* [[Server Admin Log/Archive 76|Archive 76: 2024 February]]
* [[Server Admin Log/Archive 77|Archive 77: 2024 March]]
* [[Server Admin Log/Archive 78|Archive 78: 2024 April]]
* [[Server Admin Log/Archive 79|Archive 79: 2024 May 1-15]]
* [[Server Admin Log/Archive 80|Archive 80: 2024 May 16-31]]
* [[Server Admin Log/Archive 81|Archive 81: 2024 June 1-15]]
* [[Server Admin Log/Archive 82|Archive 82: 2024 June 16-30]]
* [[Server Admin Log/Archive 83|Archive 83: 2024 July]]
* [[Server Admin Log/Archive 84|Archive 84: 2024 August]]
* [[Server Admin Log/Archive 85|Archive 85: 2024 September]]
* [[Server Admin Log/Archive 86|Archive 86: 2024 October]]
* [[Server Admin Log/Archive 87|Archive 87: 2024 November]]
* [[Server Admin Log/Archive 88|Archive 88: 2024 December]]
</div>
===2025-present===
<div style="column-count:2;-moz-column-count:2;-webkit-column-count:2">
* [[Server Admin Log/Archive 89|Archive 89: 2025 January]]
* [[Server Admin Log/Archive 90|Archive 90: 2025 February]]
* [[Server Admin Log/Archive 91|Archive 91: 2025 March]]
* [[Server Admin Log/Archive 92|Archive 92: 2025 April]]
* [[Server Admin Log/Archive 93|Archive 93: 2025 May]]
* [[Server Admin Log/Archive 94|Archive 94: 2025 June]]
* [[Server Admin Log/Archive 95|Archive 95: 2025 July]]
* [[Server Admin Log/Archive 96|Archive 96: 2025 August]]
* [[Server Admin Log/Archive 97|Archive 97: 2025 September]]
* [[Server Admin Log/Archive 98|Archive 98: 2025 October]]
* [[Server Admin Log/Archive 99|Archive 99: 2025 November]]
* [[Server Admin Log/Archive 100|Archive 100: 2025 December]]
* [[Server Admin Log/Archive 101|Archive 101: 2026 January]]
* [[Server Admin Log/Archive 102|Archive 102: 2026 February]]
* [[Server Admin Log/Archive 103|Archive 103: 2026 March]]
* [[Server Admin Log/Archive 104|Archive 104: 2026 April]]
* [[Server Admin Log/Archive 105|Archive 105: 2026 May 1-20]]
* [[Server Admin Log/Archive 106|Archive 106: 2026 May 21-June 10]]
* [[Server Admin Log/Archive 107|Archive 107: 2026 June 11-30]]
* [[Server Admin Log/Archive 108|Archive 108: 2026 July]]
* [[Server Admin Log/Archive 109|Archive 109: 2026 August]]
* [[Server Admin Log/Archive 110|Archive 110: 2026 September 1-18]]
</div>
<!-- omg! -->
<includeonly>
[[Category:Server Admin Log archive]]
</includeonly>
nbxtgagyw5e5f23dbb5w6ez8pp6r2mq
Server Admin Log
0
7919
2458699
2458694
2026-09-19T14:11:28Z
Stashbot
7414
urbanecm: Attach SHB@commonswiki to the SUL account manually (T438591, see T438591#12341750 for what I did exactly)
2458699
wikitext
text/x-wiki
== 2026-09-19 ==
* 14:11 urbanecm: Attach SHB@commonswiki to the SUL account manually ([[phab:T438591|T438591]], see [[phab:T438591|T438591]]#12341750 for what I did exactly)
* 04:08 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 04:08 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 04:08 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 04:07 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 36s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-18 ==
* 22:41 rzl@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=sessionstore,name=eqiad
* 17:08 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343065{{!}}mw.DesktopArticleTarget: if source education is enabled suppress welcome (T434249)]] (duration: 09m 26s)
* 17:05 Dreamy_Jazz: Created `securepoll_log` on `nlwiki` main DB cluster for [[phab:T434045|T434045]]
* 17:04 kemayo@deploy1003: kemayo: Continuing with deployment
* 17:03 kemayo@deploy1003: kemayo: Backport for [[gerrit:1343065{{!}}mw.DesktopArticleTarget: if source education is enabled suppress welcome (T434249)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:59 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1343065{{!}}mw.DesktopArticleTarget: if source education is enabled suppress welcome (T434249)]]
* 16:49 oblivian@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343127{{!}}ResourceLoader: hotfix for current logspam over the weekend (T438387)]] (duration: 11m 36s)
* 16:42 oblivian@deploy1003: oblivian: Continuing with deployment
* 16:42 oblivian@deploy1003: oblivian: Backport for [[gerrit:1343127{{!}}ResourceLoader: hotfix for current logspam over the weekend (T438387)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:37 oblivian@deploy1003: Started scap sync-world: Backport for [[gerrit:1343127{{!}}ResourceLoader: hotfix for current logspam over the weekend (T438387)]]
* 16:09 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ml-serve1016.eqiad.wmnet with OS trixie
* 14:49 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie
* 13:37 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host registry1004.eqiad.wmnet with OS trixie
* 13:23 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on registry1004.eqiad.wmnet with reason: host reimage
* 13:18 elukey@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on registry1004.eqiad.wmnet with reason: host reimage
* 13:04 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host registry1004.eqiad.wmnet with OS trixie
* 12:14 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:13 jclark@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:30 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 09:30 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 09:22 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on an-redacteddb1001.eqiad.wmnet with reason: cloning
* 09:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 09:19 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 09:18 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 21 hosts with reason: cloning db1270
* 09:18 marostegui: clone db1270:x4 from db1155:x4 lag will appear on x4
* 09:09 brouberol@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 09:08 brouberol@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 08:23 brouberol@dns1004: END - running authdns-update
* 08:21 brouberol@dns1004: START - running authdns-update
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 57s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-17 ==
* 21:04 tsev@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T438395|T438395]]
* 20:55 tsev@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T438395|T438395]]
* 20:47 jhuneidi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342774{{!}}Worklist Promotion test kitchen - Enable flag in production (T434513)]], [[gerrit:1342781{{!}}Exclude returntoapp query from app interception on iOS (T438395)]], [[gerrit:1342798{{!}}Revert "Update wikimania wordmark for 2026"]] (duration: 35m 59s)
* 20:35 jhuneidi@deploy1003: robertsky, jhuneidi, cmelo, tsev: Continuing with deployment
* 20:31 jhuneidi@deploy1003: robertsky, jhuneidi, cmelo, tsev: Backport for [[gerrit:1342774{{!}}Worklist Promotion test kitchen - Enable flag in production (T434513)]], [[gerrit:1342781{{!}}Exclude returntoapp query from app interception on iOS (T438395)]], [[gerrit:1342798{{!}}Revert "Update wikimania wordmark for 2026"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:11 jhuneidi@deploy1003: Started scap sync-world: Backport for [[gerrit:1342774{{!}}Worklist Promotion test kitchen - Enable flag in production (T434513)]], [[gerrit:1342781{{!}}Exclude returntoapp query from app interception on iOS (T438395)]], [[gerrit:1342798{{!}}Revert "Update wikimania wordmark for 2026"]]
* 19:15 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1100.eqiad.wmnet with OS trixie
* 19:15 vriley@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1004"
* 19:10 vriley@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1004"
* 18:52 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1100.eqiad.wmnet with reason: host reimage
* 18:48 vriley@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1100.eqiad.wmnet with reason: host reimage
* 18:32 vriley@cumin1004: START - Cookbook sre.hosts.reimage for host ms-be1100.eqiad.wmnet with OS trixie
* 18:20 urbanecm: Deploy a security fix for [[phab:T438389|T438389]]
* 17:47 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be1100.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:38 vriley@cumin1004: START - Cookbook sre.hosts.provision for host ms-be1100.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:38 vriley@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1100.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:37 vriley@cumin1004: START - Cookbook sre.hosts.provision for host ms-be1100.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:23 vriley@cumin1004: START - Cookbook sre.hosts.reimage for host ms-be1100.eqiad.wmnet with OS trixie
* 16:52 aokoth@deploy1003: Finished deploy [phabricator/deployment@c386249]: Deploy Phab (duration: 00m 12s)
* 16:52 aokoth@deploy1003: Started deploy [phabricator/deployment@c386249]: Deploy Phab
* 16:42 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1099.eqiad.wmnet with OS trixie
* 16:42 vriley@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1004"
* 16:42 vriley@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1004"
* 16:35 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be1100.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:31 aokoth@deploy1003: Finished deploy [phabricator/deployment@c386249]: Deploy Phab (duration: 00m 19s)
* 16:31 aokoth@deploy1003: Started deploy [phabricator/deployment@c386249]: Deploy Phab
* 16:21 vriley@cumin1004: START - Cookbook sre.hosts.provision for host ms-be1100.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:20 vriley@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be1100
* 16:20 vriley@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host ms-be1100
* 16:19 vriley@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:19 vriley@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [ms-be1100] - vriley@cumin1004"
* 16:19 vriley@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [ms-be1100] - vriley@cumin1004"
* 16:15 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 16:15 vriley@cumin1004: START - Cookbook sre.dns.netbox
* 16:15 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 16:14 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 16:14 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 16:12 moritzm: installing libapache-mod-auth-oidc security updates
* 16:12 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1099.eqiad.wmnet with reason: host reimage
* 16:11 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 16:11 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 16:08 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/editcheck-headless: apply
* 16:07 rzl@deploy1003: helmfile [staging] START helmfile.d/services/editcheck-headless: apply
* 16:06 vriley@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1099.eqiad.wmnet with reason: host reimage
* 16:01 moritzm: installing aom security updates
* 16:01 btullis@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ceph-admin2001.codfw.wmnet
* 16:01 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ceph-admin2001.codfw.wmnet with OS bookworm
* 15:55 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 15:55 rzl@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 15:54 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 15:54 rzl@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 15:54 rzl@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 15:53 rzl@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 15:53 rzl@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 15:52 rzl@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 15:51 vriley@cumin1004: START - Cookbook sre.hosts.reimage for host ms-be1099.eqiad.wmnet with OS trixie
* 15:48 moritzm: installing libde265 security updates
* 15:44 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ceph-admin2001.codfw.wmnet with reason: host reimage
* 15:40 btullis@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ceph-admin1001.eqiad.wmnet
* 15:40 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ceph-admin1001.eqiad.wmnet with OS bookworm
* 15:39 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir4003.*
* 15:37 btullis@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ceph-admin2001.codfw.wmnet with reason: host reimage
* 15:26 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir4003.ulsfo.wmnet with OS trixie
* 15:25 aokoth@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host phab2003.codfw.wmnet with OS trixie
* 15:23 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ceph-admin1001.eqiad.wmnet with reason: host reimage
* 15:19 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host registry1005.eqiad.wmnet with OS trixie
* 15:17 btullis@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ceph-admin1001.eqiad.wmnet with reason: host reimage
* 15:16 btullis@cumin1004: START - Cookbook sre.hosts.reimage for host ceph-admin2001.codfw.wmnet with OS bookworm
* 15:16 btullis@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ceph-admin2001.codfw.wmnet - btullis@cumin1004"
* 15:16 btullis@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ceph-admin2001.codfw.wmnet - btullis@cumin1004"
* 15:15 btullis@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ceph-admin2001.codfw.wmnet on all recursors
* 15:15 btullis@cumin1004: START - Cookbook sre.dns.wipe-cache ceph-admin2001.codfw.wmnet on all recursors
* 15:15 btullis@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:15 btullis@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ceph-admin2001.codfw.wmnet - btullis@cumin1004"
* 15:15 btullis@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ceph-admin2001.codfw.wmnet - btullis@cumin1004"
* 15:08 aokoth@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on phab2003.codfw.wmnet with reason: host reimage
* 15:06 Msz2001: Deployed private code changes to Suggestedinvestigations
* 15:05 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir4003.ulsfo.wmnet with reason: host reimage
* 15:05 aokoth@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on phab2003.codfw.wmnet with reason: host reimage
* 15:04 btullis@cumin1004: START - Cookbook sre.hosts.reimage for host ceph-admin1001.eqiad.wmnet with OS bookworm
* 15:04 btullis@cumin1004: START - Cookbook sre.dns.netbox
* 15:04 btullis@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ceph-admin1001.eqiad.wmnet - btullis@cumin1004"
* 15:04 btullis@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ceph-admin1001.eqiad.wmnet - btullis@cumin1004"
* 15:04 btullis@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ceph-admin1001.eqiad.wmnet on all recursors
* 15:04 btullis@cumin1004: START - Cookbook sre.dns.wipe-cache ceph-admin1001.eqiad.wmnet on all recursors
* 15:04 btullis@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:04 btullis@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ceph-admin1001.eqiad.wmnet - btullis@cumin1004"
* 15:04 btullis@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ceph-admin1001.eqiad.wmnet - btullis@cumin1004"
* 15:02 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on registry1005.eqiad.wmnet with reason: host reimage
* 15:01 btullis@cumin1004: START - Cookbook sre.ganeti.makevm for new host ceph-admin2001.codfw.wmnet
* 15:00 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342265{{!}}JsonSchemaBuilder: Cache the root schema in the process (T437588)]], [[gerrit:1342264{{!}}JsonSchemaBuilder: Cache the root schema in the process (T437588)]], [[gerrit:1342684{{!}}SI: Preserve the username filter when switching queues (T438308)]] (duration: 12m 34s)
* 15:00 btullis@cumin1004: START - Cookbook sre.dns.netbox
* 15:00 btullis@cumin1004: START - Cookbook sre.ganeti.makevm for new host ceph-admin1001.eqiad.wmnet
* 14:58 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir4003.ulsfo.wmnet with reason: host reimage
* 14:58 elukey@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on registry1005.eqiad.wmnet with reason: host reimage
* 14:55 urbanecm@deploy1003: mszwarc, urbanecm: Continuing with deployment
* 14:51 urbanecm@deploy1003: mszwarc, urbanecm: Backport for [[gerrit:1342265{{!}}JsonSchemaBuilder: Cache the root schema in the process (T437588)]], [[gerrit:1342264{{!}}JsonSchemaBuilder: Cache the root schema in the process (T437588)]], [[gerrit:1342684{{!}}SI: Preserve the username filter when switching queues (T438308)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:51 aokoth@cumin1004: START - Cookbook sre.hosts.reimage for host phab2003.codfw.wmnet with OS trixie
* 14:50 aokoth@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on phab2003.codfw.wmnet with reason: Reimage
* 14:47 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1342265{{!}}JsonSchemaBuilder: Cache the root schema in the process (T437588)]], [[gerrit:1342264{{!}}JsonSchemaBuilder: Cache the root schema in the process (T437588)]], [[gerrit:1342684{{!}}SI: Preserve the username filter when switching queues (T438308)]]
* 14:42 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1331830{{!}}Keep Balinese Palm Leaf variants enabled on wikisource (T436398)]], [[gerrit:1340216{{!}}Turn on variant conversion for PageAssessments (T328012)]], [[gerrit:1341949{{!}}Parsoid Read Views: Enable on 61 wikiquote wikis (T437917)]] (duration: 15m 31s)
* 14:39 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host registry1005.eqiad.wmnet with OS trixie
* 14:35 cscott@deploy1003: ssastry, cscott: Continuing with deployment
* 14:33 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir4003.ulsfo.wmnet with OS trixie
* 14:32 cscott@deploy1003: ssastry, cscott: Backport for [[gerrit:1331830{{!}}Keep Balinese Palm Leaf variants enabled on wikisource (T436398)]], [[gerrit:1340216{{!}}Turn on variant conversion for PageAssessments (T328012)]], [[gerrit:1341949{{!}}Parsoid Read Views: Enable on 61 wikiquote wikis (T437917)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:26 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1331830{{!}}Keep Balinese Palm Leaf variants enabled on wikisource (T436398)]], [[gerrit:1340216{{!}}Turn on variant conversion for PageAssessments (T328012)]], [[gerrit:1341949{{!}}Parsoid Read Views: Enable on 61 wikiquote wikis (T437917)]]
* 14:20 caro@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342678{{!}}enwiki desktop VE: add education popup for switching to source editor (T434249)]], [[gerrit:1342362{{!}}Make VE the default editor on enwiki desktop (T436574)]] (duration: 33m 52s)
* 14:07 caro@deploy1003: caro: Continuing with deployment
* 14:06 caro@deploy1003: caro: Backport for [[gerrit:1342678{{!}}enwiki desktop VE: add education popup for switching to source editor (T434249)]], [[gerrit:1342362{{!}}Make VE the default editor on enwiki desktop (T436574)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:46 caro@deploy1003: Started scap sync-world: Backport for [[gerrit:1342678{{!}}enwiki desktop VE: add education popup for switching to source editor (T434249)]], [[gerrit:1342362{{!}}Make VE the default editor on enwiki desktop (T436574)]]
* 13:37 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 13:34 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:33 elukey@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:31 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:30 elukey@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:29 Emperor: apus - radosgw-admin quota set --quota-scope=user --uid=docker-registry --max-size=5T [[phab:T438339|T438339]]
* 13:24 jclark@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1016
* 13:24 jclark@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1016
* 13:11 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:11 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:09 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:09 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:55 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:53 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 12:42 jclark@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:40 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:40 jclark@cumin1003: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:35 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 12:19 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342626{{!}}MathMathML: Simplify Mathoid fallback/a11y class logic (T436026)]], [[gerrit:1342627{{!}}ext.math.mathjax: Implement mwe-math-mathml-a11y for client-side MathJax (T436026)]] (duration: 13m 04s)
* 12:14 krinkle@deploy1003: krinkle: Continuing with deployment
* 12:10 krinkle@deploy1003: krinkle: Backport for [[gerrit:1342626{{!}}MathMathML: Simplify Mathoid fallback/a11y class logic (T436026)]], [[gerrit:1342627{{!}}ext.math.mathjax: Implement mwe-math-mathml-a11y for client-side MathJax (T436026)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:05 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1342626{{!}}MathMathML: Simplify Mathoid fallback/a11y class logic (T436026)]], [[gerrit:1342627{{!}}ext.math.mathjax: Implement mwe-math-mathml-a11y for client-side MathJax (T436026)]]
* 10:37 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 10:28 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 10:10 blake@deploy1003: Finished scap sync-world: cleanup for [[phab:T417800|T417800]] (duration: 03m 57s)
* 10:07 blake@deploy1003: Started scap sync-world: cleanup for [[phab:T417800|T417800]]
* 09:52 marostegui@cumin1004: dbctl commit (dc=all): 'Fix weights [[phab:T436496|T436496]]', diff saved to https://phabricator.wikimedia.org/P96466 and previous config saved to /var/cache/conftool/dbconfig/20260917-095235-marostegui.json
* 09:51 marostegui@cumin1004: dbctl commit (dc=all): 'Fix weights [[phab:T436496|T436496]]', diff saved to https://phabricator.wikimedia.org/P96465 and previous config saved to /var/cache/conftool/dbconfig/20260917-095131-marostegui.json
* 09:41 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:41 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 09:41 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:40 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 09:40 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:40 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 09:35 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 09:35 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 08:37 moritzm: pruned obsolete Bullseye image prometheus-nutcracker-exporter from the docker registry [[phab:T416452|T416452]]
* 08:34 XioNoX: Manually install gnmic 0.49.0 on netflow2005 - [[phab:T438291|T438291]]
* 08:28 brouberol@dns1004: END - running authdns-update
* 08:26 brouberol@dns1004: START - running authdns-update
* 08:13 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.20 refs [[phab:T430839|T430839]]
* 08:10 moritzm: imported nodejs_26.8.2-1nodesource1 to thirdparty/node26 for trixie-wikimedia [[phab:T437510|T437510]]
* 08:07 Amir1: dropped links tables from db2206 ([[phab:T437278|T437278]])
* 08:03 Amir1: dropped links tables from db2219 ([[phab:T437278|T437278]])
* 08:01 Amir1: dropped links tables from db2236 ([[phab:T437278|T437278]])
* 07:59 Amir1: dropped non-links tables from db1262 ([[phab:T437278|T437278]])
* 07:57 Amir1: dropped non-links tables from db2245 ([[phab:T437278|T437278]])
* 07:52 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@8eaca67] (releasing): [[phab:T438205|T438205]] to prod host (duration: 00m 44s)
* 07:52 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@8eaca67] (releasing): [[phab:T438205|T438205]] to prod host
* 07:49 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@8eaca67] (releasing): [[phab:T438205|T438205]] to backup host (duration: 00m 47s)
* 07:48 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@8eaca67] (releasing): [[phab:T438205|T438205]] to backup host
* 07:25 XioNoX: Manually install gnmic 0.49.0 on netflow1004 - [[phab:T438291|T438291]]
* 07:24 mlitn@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342398{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342396{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342395{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342545{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342546{{!}}Localisation updates from https://translatewiki.net.]],
* 07:19 mlitn@deploy1003: mlitn, jdlrobson: Continuing with deployment
* {{safesubst:SAL entry|1=07:18 mlitn@deploy1003: mlitn, jdlrobson: Backport for [[gerrit:1342398{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342396{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342395{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342545{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342546{{!}}Localisation updates from https://translatewiki.net.]], [[gerri}}
* 07:11 mlitn@deploy1003: Started scap sync-world: Backport for [[gerrit:1342398{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342396{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342395{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342545{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342546{{!}}Localisation updates from https://translatewiki.net.]],
* 07:10 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Jmoore111 out of all services on: 2444 hosts
* 06:07 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 05:53 marostegui@cumin1004: END (FAIL) - Cookbook sre.mysql.decommission (exit_code=99)
* 05:53 marostegui@cumin1004: Removing db1180 from zarcillo [[phab:T437222|T437222]]
* 05:53 marostegui@cumin1004: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1180.eqiad.wmnet
* 05:53 marostegui@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 05:53 marostegui@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1180.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1004"
* 05:53 marostegui@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1180.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1004"
* 05:49 marostegui@cumin1004: START - Cookbook sre.dns.netbox
* 05:44 marostegui@cumin1004: START - Cookbook sre.hosts.decommission for hosts db1180.eqiad.wmnet
* 05:43 marostegui@cumin1004: START - Cookbook sre.mysql.decommission
* 04:26 aokoth@cumin1004: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet
* 04:24 aokoth@cumin1004: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet
== 2026-09-16 ==
* 23:10 rzl: rzl@deploy1003 Finished scap sync-world: Backport for [[gerrit:1342091{{!}}Repool poolcounter[1007,2006] (T435163)]] (duration: 11m 09s)
* 22:50 rzl@deploy1003: Started scap sync-world: Backport for [[gerrit:1342091{{!}}Repool poolcounter[1007,2006] (T435163)]]
* 22:47 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342383{{!}}DonorIdentification: Confirm before unlinking donor status in preferences (T436698)]], [[gerrit:1342385{{!}}Make learn more link to new window (T438252)]] (duration: 35m 21s)
* 22:35 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 22:33 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1342383{{!}}DonorIdentification: Confirm before unlinking donor status in preferences (T436698)]], [[gerrit:1342385{{!}}Make learn more link to new window (T438252)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:12 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1342383{{!}}DonorIdentification: Confirm before unlinking donor status in preferences (T436698)]], [[gerrit:1342385{{!}}Make learn more link to new window (T438252)]]
* 22:10 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2006.codfw.wmnet
* 22:06 rzl@cumin2003: START - Cookbook sre.hosts.reboot-single for host poolcounter2006.codfw.wmnet
* 22:06 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1007.eqiad.wmnet
* 22:02 rzl@cumin2003: START - Cookbook sre.hosts.reboot-single for host poolcounter1007.eqiad.wmnet
* 21:56 rzl@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342090{{!}}Repool poolcounter[1006,2005]; depool poolcounter[1007,2006] for reboot (T435163)]] (duration: 09m 39s)
* 21:52 rzl@deploy1003: rzl: Continuing with deployment
* 21:51 rzl@deploy1003: rzl: Backport for [[gerrit:1342090{{!}}Repool poolcounter[1006,2005]; depool poolcounter[1007,2006] for reboot (T435163)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:47 rzl@deploy1003: Started scap sync-world: Backport for [[gerrit:1342090{{!}}Repool poolcounter[1006,2005]; depool poolcounter[1007,2006] for reboot (T435163)]]
* 21:46 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:43 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2005.codfw.wmnet
* 21:42 vriley@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1099.eqiad.wmnet with OS trixie
* 21:41 vriley@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1098.eqiad.wmnet with OS trixie
* 21:41 vriley@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin2003"
* 21:40 vriley@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin2003"
* 21:40 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:40 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:39 rzl@cumin2003: START - Cookbook sre.hosts.reboot-single for host poolcounter2005.codfw.wmnet
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1006.eqiad.wmnet
* 21:38 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:38 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:37 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:35 rzl@cumin2003: START - Cookbook sre.hosts.reboot-single for host poolcounter1006.eqiad.wmnet
* 21:32 rzl@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342089{{!}}Depool poolcounter[1006,2005] for reboot (T435163)]] (duration: 13m 53s)
* 21:26 rzl@deploy1003: rzl: Continuing with deployment
* 21:25 rzl@deploy1003: rzl: Backport for [[gerrit:1342089{{!}}Depool poolcounter[1006,2005] for reboot (T435163)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:23 vriley@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1098.eqiad.wmnet with reason: host reimage
* 21:18 rzl@deploy1003: Started scap sync-world: Backport for [[gerrit:1342089{{!}}Depool poolcounter[1006,2005] for reboot (T435163)]]
* 21:17 vriley@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1098.eqiad.wmnet with reason: host reimage
* 21:10 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:09 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:09 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:09 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:08 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:02 vriley@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS trixie
* 20:49 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342339{{!}}Reapply "Tell VisualEditor about the app web edit tags", modified]] (duration: 35m 55s)
* 20:37 kemayo@deploy1003: cklimas, kemayo: Continuing with deployment
* 20:33 kemayo@deploy1003: cklimas, kemayo: Backport for [[gerrit:1342339{{!}}Reapply "Tell VisualEditor about the app web edit tags", modified]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:13 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1342339{{!}}Reapply "Tell VisualEditor about the app web edit tags", modified]]
* 19:20 dwisehaupt@dns1005: END - running authdns-update
* 19:18 dwisehaupt@dns1005: START - running authdns-update
* 19:06 vriley@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:57 dwisehaupt@dns1005: END - running authdns-update
* 18:55 vriley@cumin2003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:55 dwisehaupt@dns1005: START - running authdns-update
* 18:44 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 18:42 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 18:37 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 18:35 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 18:26 robh@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:22 robh@cumin2003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:21 dzahn@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 18:21 robh@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:21 robh@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:20 dzahn@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 18:20 dzahn@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 18:18 dzahn@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 18:18 mutante: k8s/miscweb: admin_ng deploy: creating namespace for attribution.wikimedia.org [[phab:T437635|T437635]]
* 18:17 dzahn@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 18:17 dzahn@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 18:17 dzahn@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 18:16 dzahn@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 18:13 cdanis@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "fix known-client creation - cdanis@cumin1003"
* 18:13 cdanis@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: fix known-client creation - cdanis@cumin1003
* 18:12 cdanis@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: fix known-client creation - cdanis@cumin1003
* 18:12 cdanis@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "fix known-client creation - cdanis@cumin1003"
* 18:04 vriley@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:51 vriley@cumin2003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:46 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 17:46 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 17:45 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:45 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:44 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:44 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:35 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1006.eqiad.wmnet with OS trixie
* 17:35 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1003"
* 17:30 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:29 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1003"
* 17:27 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:27 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:23 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:23 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:20 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 17:20 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 17:14 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:13 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1006.eqiad.wmnet with reason: host reimage
* 17:10 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:10 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:10 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:10 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:09 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:07 jclark@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1006.eqiad.wmnet with reason: host reimage
* 17:07 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 17:05 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:05 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:52 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie
* 16:46 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir7003.magru.wmnet
* 16:44 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir7003
* 16:19 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie
* 15:42 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply
* 15:42 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply
* 15:42 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 15:42 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir7003.magru.wmnet with OS trixie
* 15:41 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 15:36 jnuche@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342279{{!}}Use parser output value instead of status (T438154)]] (duration: 33m 21s)
* 15:35 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply
* 15:35 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply
* 15:24 jnuche@deploy1003: jnuche, jforrester: Continuing with deployment
* 15:23 jnuche@deploy1003: jnuche, jforrester: Backport for [[gerrit:1342279{{!}}Use parser output value instead of status (T438154)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:19 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie
* 15:18 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir7003.magru.wmnet with reason: host reimage
* 15:14 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir7003.magru.wmnet with reason: host reimage
* 15:03 jnuche@deploy1003: Started scap sync-world: Backport for [[gerrit:1342279{{!}}Use parser output value instead of status (T438154)]]
* 14:50 moritzm: installing apache2 security updates
* 14:49 slyngshede@cumin1003: conftool action : set/pooled=yes; selector: name=cp5026.eqsin.wmnet
* 14:47 slyngshede@cumin1003: conftool action : set/weight=1; selector: name=cp5026.eqsin.wmnet
* 14:45 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5026.eqsin.wmnet with OS trixie
* 14:44 moritzm: installing python-filelock security updates
* 14:42 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir7003.magru.wmnet with OS trixie
* 14:35 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:35 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:34 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:33 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp6002.drmrs.wmnet
* 14:32 sukhe@puppetserver1001: conftool action : set/weight=100; selector: name=cp6002.drmrs.wmnet,service=ats-be
* 14:32 sukhe@puppetserver1001: conftool action : set/weight=1; selector: name=cp6002.drmrs.wmnet,service=cdn
* 14:29 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6002.drmrs.wmnet with OS trixie
* 14:24 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: sync
* 14:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:24 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:24 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: sync
* 14:14 jforrester@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.19,1.47.0-wmf.20,next --multiversion-image-basename docker-registry.discovery.wmnet/restricte
* 14:14 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:14 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:13 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1342008{{!}}abstractwiki: Add three new articles per community advice to show off the feature (T434227)]]
* 14:10 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: sync
* 14:10 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: sync
* 14:10 jforrester@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.19,1.47.0-wmf.20,next --multiversion-image-basename docker-registry.discovery.wmnet/restricte
* 14:10 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: sync
* 14:10 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: sync
* 14:09 Amir1: dropped links tables on db2237 ([[phab:T437278|T437278]])
* 14:08 Amir1: dropped links tables on db1238 ([[phab:T437278|T437278]])
* 14:07 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1342008{{!}}abstractwiki: Add three new articles per community advice to show off the feature (T434227)]]
* 14:03 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5026.eqsin.wmnet with reason: host reimage
* 14:02 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:02 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:02 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 14:01 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:00 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 13:59 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6002.drmrs.wmnet with reason: host reimage
* 13:56 slyngshede@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5026.eqsin.wmnet with reason: host reimage
* 13:54 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cp6002.drmrs.wmnet with reason: host reimage
* 13:53 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 13:52 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 13:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dse-k8s-etcd[1001-1003].eqiad.wmnet
* 13:48 bking@cumin2003: START - Cookbook sre.hosts.remove-downtime for dse-k8s-etcd[1001-1003].eqiad.wmnet
* 13:46 bking@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM dse-k8s-etcd1001.eqiad.wmnet
* 13:46 bking@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM dse-k8s-etcd1001.eqiad.wmnet
* 13:45 bking@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM dse-k8s-etcd1002.eqiad.wmnet
* 13:41 bking@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM dse-k8s-etcd1002.eqiad.wmnet
* 13:41 bking@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM dse-k8s-etcd1003.eqiad.wmnet
* 13:38 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp6002.drmrs.wmnet with OS trixie
* 13:37 bking@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM dse-k8s-etcd1003.eqiad.wmnet
* 13:37 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.reboot-vm (exit_code=99) for VM dse-k8s-etcd1003.eqiad.wmnet
* 13:37 bking@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM dse-k8s-etcd1003.eqiad.wmnet
* 13:34 slyngshede@cumin1003: START - Cookbook sre.hosts.reimage for host cp5026.eqsin.wmnet with OS trixie
* 13:34 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 13:33 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5026.mgmt.eqsin.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:29 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp6002.mgmt.drmrs.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:25 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on dse-k8s-etcd[1001-1003].eqiad.wmnet with reason: Maintenance to increase vCPUS [[phab:T438084|T438084]]
* 13:24 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:24 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:22 slyngshede@cumin1003: START - Cookbook sre.hosts.provision for host cp5026.mgmt.eqsin.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342231{{!}}SI: Unset all filters on links to cases (T434530)]], [[gerrit:1342234{{!}}SI: Unset all filters on links to cases (T434530)]] (duration: 13m 10s)
* 13:19 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org
* 13:19 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org
* 13:19 sukhe@cumin1004: START - Cookbook sre.hosts.provision for host cp6002.mgmt.drmrs.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:18 stran@deploy1003: stran: Continuing with deployment
* 13:15 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: apply
* 13:15 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: apply
* 13:13 stran@deploy1003: stran: Backport for [[gerrit:1342231{{!}}SI: Unset all filters on links to cases (T434530)]], [[gerrit:1342234{{!}}SI: Unset all filters on links to cases (T434530)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:11 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp5026.eqsin.wmnet
* 13:09 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1342231{{!}}SI: Unset all filters on links to cases (T434530)]], [[gerrit:1342234{{!}}SI: Unset all filters on links to cases (T434530)]]
* 13:09 sukhe@cumin1004: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp6002.drmrs.wmnet
* 13:05 sukhe@cumin1004: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp6002.drmrs.wmnet
* 13:05 sukhe@cumin1004: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts cp6002.drmrs.wmnet
* 13:00 dkertesz@cumin1004: conftool action : set/pooled=yes; selector: name=cp5025.eqsin.wmnet
* 12:59 dkertesz@cumin1004: conftool action : set/weight=1; selector: name=cp5025.eqsin.wmnet
* 12:52 sukhe@cumin1004: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp6002.drmrs.wmnet
* 12:52 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cp6002.mgmt.drmrs.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:51 dkertesz: eqsin pooled again ([[phab:T438052|T438052]])
* 12:49 dkertesz@cumin1004: conftool action : set/pooled=yes; selector: cluster=dnsbox,dc=eqsin
* 12:47 dkertesz@dns1004: END - running authdns-update
* 12:45 dkertesz@dns1004: START - running authdns-update
* 12:43 dkertesz@cumin1004: conftool action : set/pooled=yes; selector: cluster=dnsbox,dc=eqsin,service=authdns-update
* 12:41 dkertesz@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqsin [reason: no reason specified, [[phab:T438052|T438052]]]
* 12:41 dkertesz@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool eqsin [reason: no reason specified, [[phab:T438052|T438052]]]
* 12:38 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a1-eqiad
* 12:38 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a1-eqiad
* 12:34 sukhe@cumin1004: START - Cookbook sre.hosts.provision for host cp6002.mgmt.drmrs.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:34 sukhe@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cp6002.drmrs.wmnet with reason: reimage
* 12:33 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp6002.drmrs.wmnet
* 12:13 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5025.eqsin.wmnet with OS trixie
* 12:12 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 12:11 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-a1-eqiad
* 12:09 cmooney@cumin1004: START - Cookbook sre.network.tls for network device ssw1-a1-eqiad
* 12:01 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.20 refs [[phab:T430839|T430839]]
* 11:59 moritzm: pruned obsolete Bullseye image python3-bullseye from the docker registry [[phab:T416452|T416452]]
* 11:50 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341284{{!}}IS/IS-labs: Set wmgUseModeratorToolkit default false (T431000)]] (duration: 10m 32s)
* 11:46 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ml-lab1002.eqiad.wmnet
* 11:45 samtar@deploy1003: samtar: Continuing with deployment
* 11:44 samtar@deploy1003: samtar: Backport for [[gerrit:1341284{{!}}IS/IS-labs: Set wmgUseModeratorToolkit default false (T431000)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:41 klausman@cumin1003: START - Cookbook sre.hosts.reboot-single for host ml-lab1002.eqiad.wmnet
* 11:39 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1341284{{!}}IS/IS-labs: Set wmgUseModeratorToolkit default false (T431000)]]
* 11:39 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5025.eqsin.wmnet with reason: host reimage
* 11:35 slyngshede@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5025.eqsin.wmnet with reason: host reimage
* 11:34 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:33 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:31 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:31 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:27 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:27 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:26 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:25 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:24 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.20 refs [[phab:T430839|T430839]]
* 11:21 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply
* 11:21 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply
* 11:20 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply
* 11:20 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply
* 11:19 moritzm: kicked off a new run of production-images-weekly-rebuild.service on build2004 (previously some leftovers of buster in the config prevented a complete run)
* 11:17 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
* 11:16 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
* 11:11 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
* 11:11 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
* 11:10 jnuche@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342210{{!}}Revert "Tell VisualEditor about the app web edit tags" (T437736 T438125)]] (duration: 33m 14s)
* 11:10 slyngshede@cumin1003: START - Cookbook sre.hosts.reimage for host cp5025.eqsin.wmnet with OS trixie
* 11:05 marostegui@cumin1004: dbctl commit (dc=all): 'Remove db1180 from dbctl [[phab:T437222|T437222]]', diff saved to https://phabricator.wikimedia.org/P96459 and previous config saved to /var/cache/conftool/dbconfig/20260916-110502-marostegui.json
* 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 11:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:01 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 11:01 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 10:57 jnuche@deploy1003: jnuche: Continuing with deployment
* 10:57 jnuche@deploy1003: jnuche: Backport for [[gerrit:1342210{{!}}Revert "Tell VisualEditor about the app web edit tags" (T437736 T438125)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:37 jnuche@deploy1003: Started scap sync-world: Backport for [[gerrit:1342210{{!}}Revert "Tell VisualEditor about the app web edit tags" (T437736 T438125)]]
* 10:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5025.mgmt.eqsin.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:10 slyngshede@cumin1003: START - Cookbook sre.hosts.provision for host cp5025.mgmt.eqsin.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:02 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 10:02 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 09:58 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 09:49 slyngshede@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on cp5025.eqsin.wmnet with reason: reimaging
* 09:48 slyngshede@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cp5025.eqsin.wmnet with reason: reimaging
* 09:41 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-timeline: apply
* 09:41 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-timeline: apply
* 09:38 moritzm: imported routinator 0.15.2-1trixie to thirdparty/routinator [[phab:T438122|T438122]]
* 09:30 slyngshede@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cp5025.mgmt.eqsin.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:29 slyngshede@cumin1003: START - Cookbook sre.hosts.provision for host cp5025.mgmt.eqsin.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:19 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 09:19 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp5025.eqsin.wmnet
* 09:19 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'sync'.
* 09:18 slyngshede@cumin1003: conftool action : set/pooled=yes; selector: name=cp3074.esams.wmnet
* 09:18 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp3074.esams.wmnet
* 09:15 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 09:12 elukey@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'sync'.
* 09:12 elukey@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'sync'.
* 09:11 elukey@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'sync'.
* 09:11 elukey@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'sync'.
* 09:10 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 09:10 elukey@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 09:09 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 09:09 elukey@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 08:55 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 08:54 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 08:40 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:40 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:36 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqsin [reason: depooling for maintainance, [[phab:T438052|T438052]]]
* 08:36 slyngshede@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool eqsin [reason: depooling for maintainance, [[phab:T438052|T438052]]]
* 08:35 slyngshede@cumin1003: END (FAIL) - Cookbook sre.dns.admin (exit_code=99) DNS admin: depool eqsin [reason: no reason specified, no task ID specified]
* 08:35 slyngshede@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool eqsin [reason: no reason specified, no task ID specified]
* 08:35 slyngshede@cumin1003: conftool action : set/pooled=no; selector: cluster=dnsbox,dc=eqsin
* 08:34 fabfur: start depooling eqsin ([[phab:T438052|T438052]])
* 08:24 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.20 refs [[phab:T430839|T430839]]
* 08:22 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:22 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:14 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.20 refs [[phab:T430839|T430839]]
* 08:11 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 08:11 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 08:11 jnuche@deploy1003: Rolling back deployment
* 08:10 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 08:07 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 07:59 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 07:59 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 07:58 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 07:54 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 07:34 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply
* 07:34 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply
* 07:20 mlitn@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342112{{!}}Adds an instrument for pre-image-carousel-retest (T437076)]], [[gerrit:1342113{{!}}Adds an instrument for pre-image-carousel-retest (T437076)]], [[gerrit:1342117{{!}}Set up instrument for 5-arm test (T437076)]], [[gerrit:1342118{{!}}Set up instrument for 5-arm test (T437076)]] (duration: 10m 56s)
* 07:16 mlitn@deploy1003: mlitn: Continuing with deployment
* 07:15 mlitn@deploy1003: mlitn: Backport for [[gerrit:1342112{{!}}Adds an instrument for pre-image-carousel-retest (T437076)]], [[gerrit:1342113{{!}}Adds an instrument for pre-image-carousel-retest (T437076)]], [[gerrit:1342117{{!}}Set up instrument for 5-arm test (T437076)]], [[gerrit:1342118{{!}}Set up instrument for 5-arm test (T437076)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be veri
* 07:09 mlitn@deploy1003: Started scap sync-world: Backport for [[gerrit:1342112{{!}}Adds an instrument for pre-image-carousel-retest (T437076)]], [[gerrit:1342113{{!}}Adds an instrument for pre-image-carousel-retest (T437076)]], [[gerrit:1342117{{!}}Set up instrument for 5-arm test (T437076)]], [[gerrit:1342118{{!}}Set up instrument for 5-arm test (T437076)]]
* 06:50 moritzm: installing sudo security updates
* 06:47 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply
* 06:37 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply
* 05:12 moritzm: pruned obsolete Bullseye image buildkitd from the docker registry [[phab:T416452|T416452]]
* 04:56 kart_: Updated Apertium to 2026-09-15-084320-production ([[phab:T437213|T437213]])
* 04:54 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/apertium: apply
* 04:54 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/apertium: apply
* 04:50 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/apertium: apply
* 04:49 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/apertium: apply
* 04:45 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/apertium: apply
* 04:45 kartik@deploy1003: helmfile [staging] START helmfile.d/services/apertium: apply
* 04:24 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 04:24 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 04:22 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 04:22 oblivian@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 36s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-15 ==
* 23:09 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply
* 23:08 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply
* 23:08 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply
* 23:08 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply
* 23:07 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 23:07 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 23:07 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 23:07 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 23:06 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 23:06 rzl@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 22:57 sukhe@puppetserver1001: conftool action : set/weight=1; selector: name=cp6001.drmrs.wmnet,service=cdn
* 22:50 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc-wf1002.eqiad.wmnet with OS trixie
* 22:46 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs ([[phab:T436363|T436363]])
* 22:46 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs6003.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:46 brett@cumin2003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs6003.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:46 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs6003.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:46 brett@cumin2003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs6003.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:45 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs6002.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:45 brett@cumin2003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs6002.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:45 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs6002.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:44 brett@cumin2003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs6002.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:44 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs6001.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:44 brett@cumin2003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs6001.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:43 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs6001.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:43 brett@cumin2003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs6001.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:43 brett@cumin2003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs ([[phab:T436363|T436363]])
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc-wf1002.eqiad.wmnet with reason: host reimage
* 22:33 brett@puppetserver1001: conftool action : set/weight=100; selector: name=cp6001.*
* 22:32 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp6001.*
* 22:26 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc-wf1002.eqiad.wmnet with reason: host reimage
* 22:07 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc-wf1002
* 22:07 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-wf1002
* 22:07 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host mc-wf1002
* 22:07 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc-wf1002.eqiad.wmnet 142.48.64.10.in-addr.arpa 2.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:07 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache mc-wf1002.eqiad.wmnet 142.48.64.10.in-addr.arpa 2.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:07 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:07 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc-wf1002 - rzl@cumin2003"
* 22:07 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc-wf1002 - rzl@cumin2003"
* 22:02 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:01 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host mc-wf1002
* 22:01 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host mc-wf1002.eqiad.wmnet with OS trixie
* 21:57 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6001.drmrs.wmnet with OS trixie
* 21:55 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply
* 21:55 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply
* 21:53 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply
* 21:53 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply
* 21:53 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:53 rzl@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:52 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:52 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:48 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:48 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:34 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6001.drmrs.wmnet with reason: host reimage
* 21:30 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp6001.drmrs.wmnet with reason: host reimage
* 21:20 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342049{{!}}MobileFrontend: Add app icons (T434258)]] (duration: 11m 47s)
* 21:15 jdlrobson@deploy1003: jdlrobson, cklimas: Continuing with deployment
* 21:13 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp6001.drmrs.wmnet with OS trixie
* 21:12 jdlrobson@deploy1003: jdlrobson, cklimas: Backport for [[gerrit:1342049{{!}}MobileFrontend: Add app icons (T434258)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:12 brett@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp6001.mgmt.drmrs.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:08 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1342049{{!}}MobileFrontend: Add app icons (T434258)]]
* 20:51 cdobbins@puppetserver1001: conftool action : set/weight=1; selector: name=cp2046.codfw.wmnet
* 20:51 cdobbins@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp2046.codfw.wmnet
* 20:48 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341271{{!}}Parsoid Read Views: Enable on all namespaces on wikitech (labswiki) (T437916)]] (duration: 09m 11s)
* 20:47 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp2046.codfw.wmnet with OS trixie
* 20:43 arlolra@deploy1003: ssastry, arlolra: Continuing with deployment
* 20:42 arlolra@deploy1003: ssastry, arlolra: Backport for [[gerrit:1341271{{!}}Parsoid Read Views: Enable on all namespaces on wikitech (labswiki) (T437916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:38 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1341271{{!}}Parsoid Read Views: Enable on all namespaces on wikitech (labswiki) (T437916)]]
* 20:34 brett@cumin2003: START - Cookbook sre.hosts.provision for host cp6001.mgmt.drmrs.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:30 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp6001.*
* 20:24 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp2046.codfw.wmnet with reason: host reimage
* 20:23 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339741{{!}}Enable ReaderExperiments in eswiki, jawiki, and ptwiki (T438009)]] (duration: 15m 58s)
* 20:20 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cp2046.codfw.wmnet with reason: host reimage
* 20:19 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T436363|T436363]])
* 20:19 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs4010.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:19 brett@cumin2003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs4010.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:19 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs4010.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:19 brett@cumin2003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs4010.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:18 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs4009.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:18 arlolra@deploy1003: lwatson, arlolra: Continuing with deployment
* 20:18 brett@cumin2003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs4009.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:17 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs4009.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:17 brett@cumin2003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs4009.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:17 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs4008.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:17 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp3075.esams.wmnet
* 20:17 sukhe@puppetserver1001: conftool action : set/weight=1; selector: name=cp3075.esams.wmnet
* 20:17 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp1103.eqiad.wmnet
* 20:17 brett@cumin2003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs4008.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:16 sukhe@puppetserver1001: conftool action : set/weight=1; selector: name=cp1103.eqiad.wmnet
* 20:16 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs4008.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:16 brett@cumin2003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs4008.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:16 brett@cumin2003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T436363|T436363]])
* 20:11 arlolra@deploy1003: lwatson, arlolra: Backport for [[gerrit:1339741{{!}}Enable ReaderExperiments in eswiki, jawiki, and ptwiki (T438009)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:10 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) config_reloading A:liberica-ulsfo ([[phab:T436363|T436363]])
* 20:08 brett@cumin2003: START - Cookbook sre.loadbalancer.admin config_reloading A:liberica-ulsfo ([[phab:T436363|T436363]])
* 20:08 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp1103.eqiad.wmnet with OS trixie
* 20:07 inflatador: bking@ganeti1046 sudo gnt-instance modify -B memory=4g,vcpus=4 dse-k8s-etcd100[1-3].eqiad.wmnet [[phab:T438084|T438084]]
* 20:07 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1339741{{!}}Enable ReaderExperiments in eswiki, jawiki, and ptwiki (T438009)]]
* 20:06 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp3075.esams.wmnet with OS trixie
* 20:04 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp7009.*
* 20:04 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host cp2046.codfw.wmnet with OS trixie
* 20:02 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp7009.*
* 20:02 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp7009.*
* 20:02 brett@puppetserver1001: conftool action : set/weight=1; selector: name=cp7009.*
* 20:01 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp7009.magru.wmnet with OS trixie
* 19:53 brett@puppetserver1001: conftool action : set/weight=1; selector: name=cp4045.*
* 19:53 brett@puppetserver1001: conftool action : set/weight=1; selector: name=cp4046.*
* 19:52 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4046.*
* 19:51 cdobbins@puppetserver1001: conftool action : set/pooled=no; selector: name=cp2046.codfw.wmnet
* 19:51 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp4046.ulsfo.wmnet with OS trixie
* 19:49 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4045.*
* 19:45 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp1103.eqiad.wmnet with reason: host reimage
* 19:43 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp4045.ulsfo.wmnet with OS trixie
* 19:41 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp3075.esams.wmnet with reason: host reimage
* 19:39 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cp1103.eqiad.wmnet with reason: host reimage
* 19:37 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp7009.magru.wmnet with reason: host reimage
* 19:33 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cp3075.esams.wmnet with reason: host reimage
* 19:32 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp7009.magru.wmnet with reason: host reimage
* 19:27 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp4046.ulsfo.wmnet with reason: host reimage
* 19:23 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp4046.ulsfo.wmnet with reason: host reimage
* 19:21 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp1103.eqiad.wmnet with OS trixie
* 19:19 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp1103.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:19 sukhe@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cp1103.eqiad.wmnet with reason: reimage
* 19:18 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp4045.ulsfo.wmnet with reason: host reimage
* 19:13 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 19:12 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 19:12 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 19:12 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp3075.esams.wmnet with OS trixie
* 19:12 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp4045.ulsfo.wmnet with reason: host reimage
* 19:12 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 19:11 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp3075.mgmt.esams.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:10 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp7009.magru.wmnet with OS trixie
* 19:09 brett@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp7009.mgmt.magru.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:08 sukhe@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cp3075.esams.wmnet with reason: reimaging
* 19:08 sukhe@cumin1004: START - Cookbook sre.hosts.provision for host cp1103.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:05 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp4046.ulsfo.wmnet with OS trixie
* 19:05 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore1006.eqiad.wmnet with OS bookworm
* 19:04 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3075.*
* 19:02 brett@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp4046.mgmt.ulsfo.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:01 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp1103.eqiad.wmnet
* 19:01 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp1103.eqiad.wmnet
* 19:00 sukhe@cumin1004: START - Cookbook sre.hosts.provision for host cp3075.mgmt.esams.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:58 brett@cumin2003: START - Cookbook sre.hosts.provision for host cp7009.mgmt.magru.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:57 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3075.esams.wmnet
* 18:55 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp1101.eqiad.wmnet
* 18:55 sukhe@puppetserver1001: conftool action : set/weight=1; selector: name=cp1101.eqiad.wmnet
* 18:55 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp4045.ulsfo.wmnet with OS trixie
* 18:52 brett@cumin2003: START - Cookbook sre.hosts.provision for host cp4046.mgmt.ulsfo.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:52 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp1101.eqiad.wmnet with OS trixie
* 18:45 cdobbins@puppetserver1001: conftool action : set/weight=1; selector: name=cp2044.codfw.wmnet
* 18:44 cdobbins@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp2044.codfw.wmnet
* 18:44 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore1006.eqiad.wmnet with reason: host reimage
* 18:40 brett@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp4045.mgmt.ulsfo.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:40 eevans@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore1006.eqiad.wmnet with reason: host reimage
* 18:39 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp3074.esams.wmnet
* 18:36 sukhe@cumin1004: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) config_reloading P<nowiki>{</nowiki>lvs3008.esams.wmnet<nowiki>}</nowiki> and A:liberica
* 18:36 sukhe@cumin1004: START - Cookbook sre.loadbalancer.admin config_reloading P<nowiki>{</nowiki>lvs3008.esams.wmnet<nowiki>}</nowiki> and A:liberica
* 18:33 sukhe@puppetserver1001: conftool action : set/weight=1; selector: name=cp3074.esams.wmnet
* 18:32 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp2044.codfw.wmnet with OS trixie
* 18:31 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp3074.esams.wmnet with OS trixie
* 18:30 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp1101.eqiad.wmnet with reason: host reimage
* 18:29 brett@cumin2003: START - Cookbook sre.hosts.provision for host cp4045.mgmt.ulsfo.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:26 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cp1101.eqiad.wmnet with reason: host reimage
* 18:22 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp7009.magru.wmnet with OS trixie
* 18:20 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1006.eqiad.wmnet with OS bookworm
* 18:20 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:19 eevans@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:19 eevans@cumin1004: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore1006.eqiad.wmnet
* 18:19 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore1006.eqiad.wmnet
* 18:10 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp1101.eqiad.wmnet with OS trixie
* 18:09 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp2044.codfw.wmnet with reason: host reimage
* 18:09 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp4045.ulsfo.wmnet with OS trixie
* 18:08 eevans@cumin1004: START - Cookbook sre.hosts.reboot-single for host sessionstore1006.eqiad.wmnet
* 18:07 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp3074.esams.wmnet with reason: host reimage
* 17:52 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp7009.magru.wmnet with reason: host reimage
* 17:48 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp2044.codfw.wmnet with OS trixie
* 17:43 cdobbins@puppetserver1001: conftool action : set/pooled=no; selector: name=cp2044.codfw.wmnet
* 17:36 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp3074.esams.wmnet with OS trixie
* 17:34 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp4046.ulsfo.wmnet with OS trixie
* 17:34 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp4045.ulsfo.wmnet with OS trixie
* 17:33 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3074.esams.wmnet
* 17:28 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp4046.*
* 17:28 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp4045.*
* 17:26 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp7009.magru.wmnet with OS trixie
* 17:25 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp1101.eqiad.wmnet with OS trixie
* 17:24 cdobbins@puppetserver1001: conftool action : set/pooled=no; selector: name=cp7009.magru.wmnet
* 17:23 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp1101.eqiad.wmnet
* 17:22 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:22 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org
* 17:02 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp7009.*
* 17:01 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:01 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:00 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:59 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:59 vriley@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host ms-be1099
* 16:59 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ms-be1099
* 16:59 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:56 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 16:55 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:55 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:55 vriley@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host ms-be1099
* 16:55 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ms-be1099
* 16:55 vriley@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host ms-be1099
* 16:54 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ms-be1099
* 16:54 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1098.eqiad.wmnet with OS bullseye
* 16:53 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:53 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [ms-be1099] - vriley@cumin1003"
* 16:53 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [ms-be1099] - vriley@cumin1003"
* 16:49 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 16:33 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS bullseye
* 16:21 mutante: temp disabling puppet on C:zookeeper (32 hosts) - safe deploy of https://gerrit.wikimedia.org/r/c/operations/puppet/+/1327569
* 16:04 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341890{{!}}Restore table borders for client-side MathJax (T435274)]], [[gerrit:1340558{{!}}lift IP cap for edit-a-thon /workshop (T437609 T437594 T437470)]] (duration: 24m 19s)
* 15:59 krinkle@deploy1003: anzx, krinkle: Continuing with deployment
* 15:44 krinkle@deploy1003: anzx, krinkle: Backport for [[gerrit:1341890{{!}}Restore table borders for client-side MathJax (T435274)]], [[gerrit:1340558{{!}}lift IP cap for edit-a-thon /workshop (T437609 T437594 T437470)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:40 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1341890{{!}}Restore table borders for client-side MathJax (T435274)]], [[gerrit:1340558{{!}}lift IP cap for edit-a-thon /workshop (T437609 T437594 T437470)]]
* 15:34 elukey@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 15:34 elukey@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 15:33 brennen@deploy1003: Finished deploy [phabricator/deployment@c386249]: deploy phab1005 for [[phab:T437930|T437930]] (duration: 00m 39s)
* 15:33 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS bullseye
* 15:33 brennen@deploy1003: Started deploy [phabricator/deployment@c386249]: deploy phab1005 for [[phab:T437930|T437930]]
* 15:32 elukey@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 15:32 elukey@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'sync'.
* 15:32 brennen@deploy1003: Finished deploy [phabricator/deployment@c386249]: deploy phab2003 for [[phab:T437930|T437930]] (duration: 00m 52s)
* 15:32 elukey@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 15:32 elukey@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'sync'.
* 15:31 brennen@deploy1003: Started deploy [phabricator/deployment@c386249]: deploy phab2003 for [[phab:T437930|T437930]]
* 15:31 elukey@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'.
* 15:31 elukey@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'.
* 15:26 jelto@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on phab2003.codfw.wmnet,phab[1005-1006].eqiad.wmnet with reason: Phabricator deploy
* 15:26 moritzm: pruned obsolete Bullseye image amd-gpu-tester from the docker registry [[phab:T416452|T416452]]
* 15:12 elukey@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 15:12 elukey@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 15:11 elukey@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 15:11 elukey@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 15:00 tgr@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341292{{!}}CommonSettings: Use a restrictive CSP for auth.wikimedia.org (T419684)]] (duration: 25m 11s)
* 14:55 tgr@deploy1003: tgr, arendpieter: Continuing with deployment
* 14:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:53 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:53 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:52 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:49 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 14:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 14:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 14:47 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 14:47 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 14:45 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 14:45 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 14:45 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 14:44 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 14:42 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 14:42 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 14:39 tgr@deploy1003: tgr, arendpieter: Backport for [[gerrit:1341292{{!}}CommonSettings: Use a restrictive CSP for auth.wikimedia.org (T419684)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:34 tgr@deploy1003: Started scap sync-world: Backport for [[gerrit:1341292{{!}}CommonSettings: Use a restrictive CSP for auth.wikimedia.org (T419684)]]
* 14:17 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341895{{!}}ReportIncidentController: Instance cache expensive methods (T437588)]] (duration: 11m 56s)
* 14:16 btullis@cumin1004: END (PASS) - Cookbook sre.ceph.rotate-osd-keys (exit_code=0) rolling rotate_keys on A:cephosd-codfw
* 14:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:12 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 14:09 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1341895{{!}}ReportIncidentController: Instance cache expensive methods (T437588)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:05 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1341895{{!}}ReportIncidentController: Instance cache expensive methods (T437588)]]
* 13:43 btullis@cumin1004: START - Cookbook sre.ceph.rotate-osd-keys rolling rotate_keys on A:cephosd-codfw
* 13:36 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341861{{!}}SuggestedInvestigations: Update "sockpuppet" queue view defaults (T438018)]] (duration: 10m 23s)
* 13:32 stran@deploy1003: stran: Continuing with deployment
* 13:30 stran@deploy1003: stran: Backport for [[gerrit:1341861{{!}}SuggestedInvestigations: Update "sockpuppet" queue view defaults (T438018)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:26 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1341861{{!}}SuggestedInvestigations: Update "sockpuppet" queue view defaults (T438018)]]
* 13:21 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:20 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:19 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334946{{!}}ArticleGuidance: Remove the experiment configuration keys (T434487)]] (duration: 09m 19s)
* 13:16 btullis@cumin1004: END (PASS) - Cookbook sre.ceph.rotate-osd-keys (exit_code=0) rolling rotate_keys on P<nowiki>{</nowiki>cephosd2001.codfw.wmnet<nowiki>}</nowiki> and (A:cephosd-codfw or A:cephosd-eqiad)
* 13:15 sbisson@deploy1003: sbisson: Continuing with deployment
* 13:14 sbisson@deploy1003: sbisson: Backport for [[gerrit:1334946{{!}}ArticleGuidance: Remove the experiment configuration keys (T434487)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) config_reloading P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:10 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1334946{{!}}ArticleGuidance: Remove the experiment configuration keys (T434487)]]
* 13:10 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin config_reloading P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:08 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:08 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:07 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:07 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:07 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:07 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:03 btullis@cumin1004: START - Cookbook sre.ceph.rotate-osd-keys rolling rotate_keys on P<nowiki>{</nowiki>cephosd2001.codfw.wmnet<nowiki>}</nowiki> and (A:cephosd-codfw or A:cephosd-eqiad)
* 13:00 btullis@cumin1004: END (PASS) - Cookbook sre.ceph.rotate-osd-keys (exit_code=0) rolling rotate_keys on P<nowiki>{</nowiki>cephosd2001.codfw.wmnet<nowiki>}</nowiki> and (A:cephosd-codfw or A:cephosd-eqiad)
* 12:59 btullis@cumin1004: START - Cookbook sre.ceph.rotate-osd-keys rolling rotate_keys on P<nowiki>{</nowiki>cephosd2001.codfw.wmnet<nowiki>}</nowiki> and (A:cephosd-codfw or A:cephosd-eqiad)
* 12:46 btullis@cumin1004: END (PASS) - Cookbook sre.ceph.rotate-osd-keys (exit_code=0) rolling rotate_keys on P<nowiki>{</nowiki>cephosd2001.codfw.wmnet<nowiki>}</nowiki> and (A:cephosd-codfw or A:cephosd-eqiad)
* 12:45 btullis@cumin1004: START - Cookbook sre.ceph.rotate-osd-keys rolling rotate_keys on P<nowiki>{</nowiki>cephosd2001.codfw.wmnet<nowiki>}</nowiki> and (A:cephosd-codfw or A:cephosd-eqiad)
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: network maintenance finished, [[phab:T437984|T437984]]]
* 12:40 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: network maintenance finished, [[phab:T437984|T437984]]]
* 12:29 XioNoX: asw1-b4-magru> request system reboot - [[phab:T437984|T437984]]
* 12:24 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 12:24 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 12:24 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 12:23 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 12:23 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 12:23 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 12:22 moritzm: installing shadow security updates
* 12:19 slyngshede@puppetserver1001: conftool action : set/weight=1; selector: name=cp7010.magru.wmnet
* 12:13 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 12:13 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 12 hosts with reason: Switch maintenance
* 12:12 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-b4-magru,asw1-b4-magru IPv6,asw1-b4-magru.mgmt with reason: Switch maintenance
* 12:11 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 12:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: switch reboot, [[phab:T437984|T437984]]]
* 12:11 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: switch reboot, [[phab:T437984|T437984]]]
* 12:09 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on install7002.wikimedia.org with reason: switch reboot
* 12:08 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 12:07 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 12:07 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 12:07 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 12:06 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 12:06 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 12:03 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 12:03 XioNoX: push pfw policies - [[phab:T437627|T437627]]
* 12:01 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 12:01 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) config_reloading P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 12:00 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin config_reloading P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 11:56 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 11:56 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 11:33 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2250.codfw.wmnet with reason: cloning db2201
* 11:18 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7004.magru.wmnet
* 11:17 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7004.magru.wmnet
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 11:05 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 11:05 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 11:04 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 10:51 slyngshede@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp7010.magru.wmnet
* 10:34 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 10:33 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 10:30 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1098.eqiad.wmnet with OS trixie
* 10:21 moritzm: failover Ganeti master in magru to ganeti7001
* 10:20 moritzm: increased DRBD replication speed in Ganeti/magru [[phab:T428878|T428878]]
* 10:10 hashar@deploy1003: Finished deploy [integration/docroot@5cf09c8]: build: Updating npm dependencies (duration: 00m 13s)
* 10:10 hashar@deploy1003: Started deploy [integration/docroot@5cf09c8]: build: Updating npm dependencies
* 10:09 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 10:08 moritzm: increased DRBD replication speed in Ganeti/esams [[phab:T428878|T428878]]
* 10:07 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 10:05 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 10:05 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:39 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: switches reboot, [[phab:T437984|T437984]]]
* 09:39 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool esams [reason: switches reboot, [[phab:T437984|T437984]]]
* 09:31 XioNoX: asw1-by27-esams> request system reboot - [[phab:T437984|T437984]]
* 09:30 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS trixie
* 09:28 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp7010.magru.wmnet with OS trixie
* 09:26 ayounsi@cumin1003: END (FAIL) - Cookbook sre.network.depool-rack (exit_code=99) with action 'depool' for esams rack BY27
* 09:24 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for esams rack BY27
* 09:24 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1098.eqiad.wmnet with OS trixie
* 09:23 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS trixie
* 09:22 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be1098.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 09:15 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.20 refs [[phab:T430839|T430839]]
* 09:10 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1098.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 09:06 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-by27-esams,asw1-by27-esams IPv6,asw1-by27-esams.mgmt with reason: Switch maintenance
* 09:05 ayounsi@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on asw1-by27-esams IPv6,asw1-by27-esams.mgmt,asw1-by-27-esams with reason: Switch maintenance
* 09:04 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 12 hosts with reason: Switch maintenance
* 09:04 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp7010.magru.wmnet with reason: host reimage
* 09:01 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: switches reboot, [[phab:T437984|T437984]]]
* 09:00 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: switches reboot, [[phab:T437984|T437984]]]
* 09:00 slyngshede@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp7010.magru.wmnet with reason: host reimage
* 08:59 jnuche@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341698{{!}}RestSandbox: Pass JsonLocalizer instead of ResponseFactory to ModuleManager (T437982)]] (duration: 12m 03s)
* 08:55 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2197.codfw.wmnet with reason: cloning db2201
* 08:55 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on install3004.wikimedia.org with reason: switch reboot
* 08:53 jnuche@deploy1003: jnuche: Continuing with deployment
* 08:52 jnuche@deploy1003: jnuche: Backport for [[gerrit:1341698{{!}}RestSandbox: Pass JsonLocalizer instead of ResponseFactory to ModuleManager (T437982)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:49 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: sync
* 08:49 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: sync
* 08:47 jnuche@deploy1003: Started scap sync-world: Backport for [[gerrit:1341698{{!}}RestSandbox: Pass JsonLocalizer instead of ResponseFactory to ModuleManager (T437982)]]
* 08:33 slyngshede@cumin1003: START - Cookbook sre.hosts.reimage for host cp7010.magru.wmnet with OS trixie
* 08:26 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be1098.eqiad.wmnet with OS trixie
* 08:26 slyngshede@puppetserver1001: conftool action : set/pooled=no; selector: name=cp7010.magru.wmnet
* 08:25 XioNoX: asw1-b3-magru - Disable logging and file logging for BRCM_PKT - [[phab:T437984|T437984]]
* 08:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS trixie
* 08:18 dpogorzelski@dns1004: END - running authdns-update
* 08:15 dpogorzelski@dns1004: START - running authdns-update
* 08:14 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3005.esams.wmnet
* 08:13 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3005.esams.wmnet
* 08:07 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341274{{!}}SI: Implement "queue view" functionality (T437183)]], [[gerrit:1341242{{!}}SuggestedInvestigations: Add and enable 'sockpuppets' queue view (T437183)]], [[gerrit:1341254{{!}}Add wmf-specific Special:SuggestedInvestigations messages (T437183)]] (duration: 55m 27s)
* 07:54 stran@deploy1003: stran: Continuing with deployment
* 07:31 stran@deploy1003: stran: Backport for [[gerrit:1341274{{!}}SI: Implement "queue view" functionality (T437183)]], [[gerrit:1341242{{!}}SuggestedInvestigations: Add and enable 'sockpuppets' queue view (T437183)]], [[gerrit:1341254{{!}}Add wmf-specific Special:SuggestedInvestigations messages (T437183)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:18 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply
* 07:18 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply
* 07:11 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1341274{{!}}SI: Implement "queue view" functionality (T437183)]], [[gerrit:1341242{{!}}SuggestedInvestigations: Add and enable 'sockpuppets' queue view (T437183)]], [[gerrit:1341254{{!}}Add wmf-specific Special:SuggestedInvestigations messages (T437183)]]
* 07:06 moritzm: pruned obsolete Bullseye image python3-devel from the docker registry [[phab:T416452|T416452]]
* 06:51 moritzm: pruned obsolete Bullseye image python3-build-bullseye from the docker registry [[phab:T416452|T416452]]
* 05:59 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply
* 05:49 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply
* 05:48 oblivian@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 05:47 oblivian@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 05:38 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply
* 05:28 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-video: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-video: apply
* 05:10 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-video: apply
* 05:10 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-video: apply
* 05:10 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-video: apply
* 05:10 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-video: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-syntaxhighlight: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-syntaxhighlight: apply
* 05:10 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-syntaxhighlight: apply
* 05:10 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-syntaxhighlight: apply
* 05:10 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-syntaxhighlight: apply
* 05:10 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-syntaxhighlight: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-media: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-media: apply
* 05:10 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-media: apply
* 05:10 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-media: apply
* 05:10 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-media: apply
* 05:10 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-media: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-constraints: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-constraints: apply
* 05:10 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-constraints: apply
* 05:09 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-constraints: apply
* 05:09 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-constraints: apply
* 05:09 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-constraints: apply
* 05:08 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply
* 05:08 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox: apply
* 05:07 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply
* 05:07 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox: apply
* 05:07 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox: apply
* 05:07 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox: apply
* 05:07 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply
* 05:07 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply
* 05:06 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-timeline: apply
* 05:06 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-timeline: apply
* 05:06 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply
* 05:06 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply
* 04:07 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.17 (duration: 07m 10s)
* 03:06 mwpresync@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.19,1.47.0-wmf.20,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.20 refs [[phab:T430839|T430839]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 22s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:43 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 00:23 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore1005.eqiad.wmnet with reason: host reimage
* 00:19 jclark@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore1005.eqiad.wmnet with reason: host reimage
* 00:17 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 00:07 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sessionstore1005.eqiad.wmnet with OS bookworm
== 2026-09-14 ==
* 23:41 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 23:26 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324966{{!}}Enable Produnto on pilot wikis (T421436)]] (duration: 12m 59s)
* 23:22 tstarling@deploy1003: tstarling: Continuing with deployment
* 23:17 tstarling@deploy1003: tstarling: Backport for [[gerrit:1324966{{!}}Enable Produnto on pilot wikis (T421436)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:13 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1324966{{!}}Enable Produnto on pilot wikis (T421436)]]
* 23:01 eevans@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 22:41 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341385{{!}}VisualEditor: don't register settings tool in wikitextCommandRegistry (T437810)]] (duration: 09m 22s)
* 22:36 kemayo@deploy1003: kemayo: Continuing with deployment
* 22:36 kemayo@deploy1003: kemayo: Backport for [[gerrit:1341385{{!}}VisualEditor: don't register settings tool in wikitextCommandRegistry (T437810)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:31 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1341385{{!}}VisualEditor: don't register settings tool in wikitextCommandRegistry (T437810)]]
* 22:22 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 22:22 eevans@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 21:43 sbassett: Deployed security fix for [[phab:T435623|T435623]]
* 21:29 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1098.eqiad.wmnet with OS bullseye
* 21:29 sbassett: Deployed security fix for [[phab:T434372|T434372]]
* 21:26 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 21:26 eevans@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 21:05 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338999{{!}}Enable ReadingLists for all logged-in users on English Wikipedia (T434923)]], [[gerrit:1340004{{!}}Enable Reading Recommendations experiment on test wiki (T437665)]] (duration: 11m 03s)
* 21:00 aude@deploy1003: aude, jdlrobson: Continuing with deployment
* 20:58 aude@deploy1003: aude, jdlrobson: Backport for [[gerrit:1338999{{!}}Enable ReadingLists for all logged-in users on English Wikipedia (T434923)]], [[gerrit:1340004{{!}}Enable Reading Recommendations experiment on test wiki (T437665)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:54 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1338999{{!}}Enable ReadingLists for all logged-in users on English Wikipedia (T434923)]], [[gerrit:1340004{{!}}Enable Reading Recommendations experiment on test wiki (T437665)]]
* 20:47 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341278{{!}}[A11y] Add list semantics to ReadingList page (T435864 T434923)]] (duration: 12m 49s)
* 20:43 aude@deploy1003: aude, jdlrobson: Continuing with deployment
* 20:39 aude@deploy1003: aude, jdlrobson: Backport for [[gerrit:1341278{{!}}[A11y] Add list semantics to ReadingList page (T435864 T434923)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:34 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1341278{{!}}[A11y] Add list semantics to ReadingList page (T435864 T434923)]]
* 20:32 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339811{{!}}wmf-config: Register content/v2-beta REST module as disabled (T432798)]], [[gerrit:1338274{{!}}wikifunctions: Move abstract fragments to mainstash (T432849)]] (duration: 25m 21s)
* 20:28 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:27 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:27 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:27 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:25 jforrester@deploy1003: jforrester, aghirelli: Continuing with deployment
* 20:24 jforrester@deploy1003: jforrester, aghirelli: Backport for [[gerrit:1339811{{!}}wmf-config: Register content/v2-beta REST module as disabled (T432798)]], [[gerrit:1338274{{!}}wikifunctions: Move abstract fragments to mainstash (T432849)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS bullseye
* 20:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be1098.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 20:06 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1339811{{!}}wmf-config: Register content/v2-beta REST module as disabled (T432798)]], [[gerrit:1338274{{!}}wikifunctions: Move abstract fragments to mainstash (T432849)]]
* 20:04 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1098.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 20:03 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be1098
* 20:02 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ms-be1098
* 20:02 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 20:02 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [ms-be1098] - vriley@cumin1003"
* 20:02 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [ms-be1098] - vriley@cumin1003"
* 19:59 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 19:58 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 19:57 dzahn@dns1005: END - running authdns-update
* 19:55 eevans@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 19:55 dzahn@dns1005: START - running authdns-update
* 19:54 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 19:38 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341273{{!}}[CodeMirror] enable for new users (enwiki), new VE integration (global) (T288161 T432558)]] (duration: 33m 51s)
* 19:26 musikanimal@deploy1003: musikanimal: Continuing with deployment
* 19:22 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1341273{{!}}[CodeMirror] enable for new users (enwiki), new VE integration (global) (T288161 T432558)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:04 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1341273{{!}}[CodeMirror] enable for new users (enwiki), new VE integration (global) (T288161 T432558)]]
* 18:53 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:50 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 18:50 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 18:49 eevans@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 18:26 brett@cumin2003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d6-eqiad
* 18:26 brett@cumin2003: START - Cookbook sre.network.tls for network device lsw1-d6-eqiad
* 18:26 brett@cumin2003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-eqiad
* 18:26 brett@cumin2003: START - Cookbook sre.network.tls for network device ssw1-d8-eqiad
* 18:25 brett@cumin2003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-eqiad
* 18:25 brett@cumin2003: START - Cookbook sre.network.tls for network device lsw1-d4-eqiad
* 18:25 root@cumin2003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-eqiad
* 18:25 root@cumin2003: START - Cookbook sre.network.tls for network device lsw1-d2-eqiad
* 18:17 jclark@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:14 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337611{{!}}extension-list: Add ModeratorToolkit (T431000)]] (duration: 09m 34s)
* 18:10 samtar@deploy1003: samtar: Continuing with deployment
* 18:09 samtar@deploy1003: samtar: Backport for [[gerrit:1337611{{!}}extension-list: Add ModeratorToolkit (T431000)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:06 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 18:05 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1337611{{!}}extension-list: Add ModeratorToolkit (T431000)]]
* 18:01 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 17:48 jclark@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 17:07 tgr@deploy1003: Finished scap sync-world: Backport for [[gerrit:1330446{{!}}CommonSettings: Use a restrictive, eval-free CSP for auth.wikimedia.org (T419684)]] (duration: 23m 19s)
* 17:00 tgr@deploy1003: arendpieter, tgr: Rolling back deployment
* 16:52 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 16:51 eevans@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 16:49 tgr@deploy1003: arendpieter, tgr: Backport for [[gerrit:1330446{{!}}CommonSettings: Use a restrictive, eval-free CSP for auth.wikimedia.org (T419684)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:44 tgr@deploy1003: Started scap sync-world: Backport for [[gerrit:1330446{{!}}CommonSettings: Use a restrictive, eval-free CSP for auth.wikimedia.org (T419684)]]
* 16:09 Amir1: drop links tables from db1252 ([[phab:T437278|T437278]])
* 16:07 Amir1: drop links tables from db2240 ([[phab:T437278|T437278]])
* 16:05 Amir1: drop non-links tables from db2247 ([[phab:T437278|T437278]])
* 15:53 Lucas_WMDE: UTC afternoon backport+config window belatedly done
* 15:50 lucaswerkmeister-wmde@deploy1003: mwscript-k8s job started: namespaceDupes abstractwiki --fix # [[phab:T437772|T437772]]
* 15:49 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335401{{!}}Adjust extendedconfirmed calculation to first edit on viwiki (T437006)]], [[gerrit:1340505{{!}}core-Namespaces: Add AW and AWT alias for its talk in abstractwiki (T437772)]] (duration: 10m 23s)
* 15:48 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 15:47 eevans@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 15:45 lucaswerkmeister-wmde@deploy1003: bunnypranav, lucaswerkmeister-wmde, tryvix1509: Continuing with deployment
* 15:43 lucaswerkmeister-wmde@deploy1003: bunnypranav, lucaswerkmeister-wmde, tryvix1509: Backport for [[gerrit:1335401{{!}}Adjust extendedconfirmed calculation to first edit on viwiki (T437006)]], [[gerrit:1340505{{!}}core-Namespaces: Add AW and AWT alias for its talk in abstractwiki (T437772)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:39 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1335401{{!}}Adjust extendedconfirmed calculation to first edit on viwiki (T437006)]], [[gerrit:1340505{{!}}core-Namespaces: Add AW and AWT alias for its talk in abstractwiki (T437772)]]
* 15:36 elukey@dns1004: END - running authdns-update
* 15:36 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 15:35 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 15:35 moritzm: installing shadow security updates
* 15:35 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 15:34 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:33 elukey@dns1004: START - running authdns-update
* 15:33 eevans@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:32 eevans@cumin1004: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore1005.eqiad.wmnet
* 15:29 lucaswerkmeister-wmde@deploy1003: mwscript-k8s job started: namespaceDupes afwiki --fix # [[phab:T437576|T437576]]
* 15:29 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338902{{!}}afwiki: Create Draft and Draft talk namespaces (T437576)]] (duration: 15m 44s)
* 15:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 15:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 15:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 15:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 15:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 15:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 15:21 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, tryvix1509: Continuing with deployment
* 15:21 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:17 eevans@cumin1004: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore1005.eqiad.wmnet
* 15:17 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, tryvix1509: Backport for [[gerrit:1338902{{!}}afwiki: Create Draft and Draft talk namespaces (T437576)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:17 eevans@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore1005.eqiad.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 15:13 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore1004.eqiad.wmnet
* 15:13 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1338902{{!}}afwiki: Create Draft and Draft talk namespaces (T437576)]]
* 15:06 eevans@cumin1004: START - Cookbook sre.hosts.reboot-single for host sessionstore1004.eqiad.wmnet
* 14:59 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore1004.eqiad.wmnet with OS bookworm
* 14:38 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore1004.eqiad.wmnet with reason: host reimage
* 14:33 marostegui@dns1004: END - running authdns-update
* 14:32 eevans@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore1004.eqiad.wmnet with reason: host reimage
* 14:30 marostegui@dns1004: START - running authdns-update
* 14:15 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1004.eqiad.wmnet with OS bookworm
* 14:14 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore1004.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:13 eevans@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1004.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:13 eevans@cumin1004: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore1004.eqiad.wmnet
* 14:13 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore1004.eqiad.wmnet
* 14:05 moritzm: kick off a rebuild of base images on build2004
* 14:05 moritzm: kick off a rebuild of base images on build2004
* 14:00 eevans@cumin1004: START - Cookbook sre.hosts.reboot-single for host sessionstore1004.eqiad.wmnet
* 14:00 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:00 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 13:56 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:54 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:52 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:44 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 13:43 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 13:42 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:40 eevans@cumin1004: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore1004.eqiad.wmnet
* 13:40 eevans@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore1004.eqiad.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 13:40 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:40 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:39 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:39 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:39 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 13:38 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 13:38 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:38 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:38 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 13:37 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 13:37 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:36 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:36 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 13:35 oblivian@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:35 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 13:35 oblivian@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:35 oblivian@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:34 oblivian@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'.
* 13:34 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 13:34 oblivian@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:34 oblivian@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:34 oblivian@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:34 oblivian@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 13:34 oblivian@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'.
* 13:33 oblivian@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'.
* 13:33 oblivian@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
* 13:33 oblivian@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
* 13:33 oblivian@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 13:33 oblivian@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 13:33 oblivian@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 13:32 oblivian@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 13:32 oblivian@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 13:32 oblivian@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 13:32 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 13:32 oblivian@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 13:32 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 13:32 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 13:29 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:23 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1003.eqiad.wmnet
* 13:21 sukhe: sudo cumin -b11 "A:cp-text" "run-puppet-agent --enable 'merging CR 1338134'" [[phab:T425441|T425441]]
* 13:20 sukhe: sudo cumin -b11 "A:cp-text" "run-puppet-agent --enable 'merging CR 1338134'"[[phab:T425441|T425441]]
* 13:19 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:17 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1003.eqiad.wmnet
* 13:14 moritzm: installing Bird security updates
* 13:09 sukhe: sudo cumin "A:cp-text" "disable-puppet 'merging CR 1338134'"
* 13:06 jmm@dns1004: END - running authdns-update
* 13:04 jmm@dns1004: START - running authdns-update
* 12:58 moritzm: update Trixie installer image to 13.7 [[phab:T437715|T437715]]
* 12:58 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 12:54 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 12:52 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 12:49 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 12:48 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 12:47 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 12:44 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 12:44 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 12:42 oblivian@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 12:40 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 12:40 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 12:40 oblivian@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 12:39 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 12:39 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 12:39 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 12:37 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 12:37 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 12:36 marostegui@cumin1004: conftool action : set/pooled=yes; selector: name=clouddb1025.eqiad.wmnet,service=x4
* 12:34 _joe_: adding gvisor labels to all wikikube clusters nodes
* 12:30 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 12:14 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 12:14 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 11:22 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1260: After cloning
* 10:48 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:45 ladsgroup@dns1004: END - running authdns-update
* 10:42 ladsgroup@dns1004: START - running authdns-update
* 10:37 marostegui@cumin1004: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=x4
* 10:37 marostegui@cumin1004: START - Cookbook sre.mysql.pool pool db1260: After cloning
* 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 10:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 10:04 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 10:04 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 09:53 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:53 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:41 Amir1: drop links tables from db2172 ([[phab:T437278|T437278]])
* 09:40 Amir1: drop links tables from db1228 ([[phab:T437278|T437278]])
* 09:08 marostegui: Stop mariadb on db1260 to clone dbstore1007, there will be lag on wikireplicas:x4 https://phabricator.wikimedia.org/T437839
* 09:07 taavi@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1025.eqiad.wmnet
* 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1260: Needs to clone another host from this one
* 09:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1260: Needs to clone another host from this one
* 09:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb[1024-1025].eqiad.wmnet,db[1155,1260].eqiad.wmnet,dbstore1007.eqiad.wmnet with reason: Adding x4
* 08:44 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:42 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:41 moritzm: pruned obsolete Bullseye images php8.3-icu72-cli / php8.3-icu72-fpm-multiversion-base / php8.3-icu72-fpm from the docker registry [[phab:T416452|T416452]]
* 08:37 taavi@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet
* 08:37 taavi@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1024.eqiad.wmnet
* 08:36 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on dbstore1007.eqiad.wmnet with reason: Adding x4
* 08:35 moritzm: pruned obsolete Bullseye images php8.1-cli/php8.1-fpm/ php8.1-fpm-multiversion-base from the docker registry [[phab:T416452|T416452]]
* 08:10 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: sync
* 08:08 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: sync
* 07:58 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 23s)
* 07:57 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 07:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 07:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 07:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 07:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 07:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 07:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 07:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 07:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 07:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 07:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 07:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 07:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 07:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 07:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 07:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 07:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 07:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 07:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 07:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 07:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 07:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 07:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 07:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 07:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 07:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 07:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 07:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 07:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 07:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 07:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 07:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 07:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 07:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 07:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 07:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 07:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 07:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 07:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 07:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 07:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 07:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 07:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 07:31 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1340801{{!}}Allow title-like strings with Package: prefix in require() (T430644)]], [[gerrit:1340802{{!}}Runtime: Add a facility for loading files by title (T430644)]] (duration: 34m 30s)
* 07:18 tstarling@deploy1003: tstarling: Continuing with deployment
* 07:17 tstarling@deploy1003: tstarling: Backport for [[gerrit:1340801{{!}}Allow title-like strings with Package: prefix in require() (T430644)]], [[gerrit:1340802{{!}}Runtime: Add a facility for loading files by title (T430644)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 06:56 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1340801{{!}}Allow title-like strings with Package: prefix in require() (T430644)]], [[gerrit:1340802{{!}}Runtime: Add a facility for loading files by title (T430644)]]
* 06:26 TimStarling: on deploy1003: docker image pull docker-registry.wikimedia.org/php8.3-fpm-multiversion-base
* 05:51 _joe_: pulled bookworm:latest from build2004 to build2001 [[phab:T437829|T437829]]
* 05:39 _joe_: force-running build-base-images on build2004 for [[phab:T437829|T437829]]
* 04:53 TimStarling: on build2001 rebuilding base images [[phab:T437829|T437829]]
* 03:00 tstarling@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.18,1.47.0-wmf.19,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted
* 02:59 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1340801{{!}}Allow title-like strings with Package: prefix in require() (T430644)]], [[gerrit:1340802{{!}}Runtime: Add a facility for loading files by title (T430644)]]
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-13 ==
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 29s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-12 ==
* 19:40 ladsgroup@cumin1003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-eqiad
* 19:32 ladsgroup@cumin1003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-eqiad
* 19:30 ladsgroup@cumin1003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw
* 19:21 ladsgroup@cumin1003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 35s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-11 ==
* 21:51 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 21:50 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 16:47 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339808{{!}}Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813{{!}}Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]] (duration: 07m 23s)
* 16:43 sbassett@deploy1003: sbassett: Continuing with deployment
* 16:42 sbassett@deploy1003: sbassett: Backport for [[gerrit:1339808{{!}}Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813{{!}}Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:40 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1339808{{!}}Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813{{!}}Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]]
* 16:08 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 14:39 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 14:10 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:10 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:10 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:09 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 13:40 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:28 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:11 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:10 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: db1199 repool
* 11:05 moritzm: installing Linux 6.1.187 on Bookworm hosts
* 11:05 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Jcrespo out of all services on: 2443 hosts
* 10:44 aokoth@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2020 in turn
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2019 in turn
* 10:41 Emperor: restart versitygw@objectstorage0[0-3].service on backup2018 in turn
* 10:40 Emperor: restart versitygw@objectstorage0[0-3].service on backup2017 in turn
* 10:39 Emperor: restart versitygw@objectstorage0[0-3].service on backup2016 in turn
* 10:37 Emperor: restart versitygw@objectstorage0[0-3].service on backup2015 in turn
* 10:36 Emperor: restart versitygw@objectstorage0[0-3].service on backup1020 in turn
* 10:35 Emperor: restart versitygw@objectstorage0[0-3].service on backup1019 in turn
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1199: db1199 repool
* 10:33 Emperor: restart versitygw@objectstorage0[0-3].service on backup1018 in turn
* 10:32 Emperor: restart versitygw@objectstorage0[0-3].service on backup1017 in turn
* 10:30 Emperor: restart versitygw@objectstorage0[0-3].service on backup1016 in turn
* 10:20 Emperor: restart versitygw@objectstorage0[1-3].service on backup1015 in turn
* 10:17 Emperor: restart versitygw@objectstorage00.service on backup1015
* 10:15 aokoth@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 08:46 slyngs: Update CAS/SSO to CAS 7.3.8.3
* 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:45 slyngshede@dns1004: END - running authdns-update
* 08:43 slyngshede@dns1004: START - running authdns-update
* 08:36 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:28 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 7 hosts with reason: Restarting s5
* 08:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db[1154,1269].eqiad.wmnet with reason: Restarting s5
* 08:20 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:20 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Repooling db1159
* 07:58 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:58 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:54 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:28 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:27 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1199: Needs to clone another host from this one
* 07:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1199: Needs to clone another host from this one
* 07:16 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Repooling db1159
* 07:15 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1199.eqiad.wmnet with reason: Cloning s4
* 07:10 TimStarling: killed jobs for [[phab:T437056|T437056]] since they weren't purging
* 06:38 TimStarling: also started refreshLinks for ptwiki and zhwiki, reparsing ~3000 pages altogether [[phab:T437056|T437056]]
* 06:27 TimStarling: for [[phab:T437056|T437056]]: mwscript-k8s refreshLinks.php --wiki=eswiki --tracking-category scribunto-common-error-category
* 05:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1159: Needs to clone another host from this one
* 05:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1159: Needs to clone another host from this one
* 05:30 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1159.eqiad.wmnet with reason: Cloning
* 05:29 marostegui: Start cloning db1245:s5 [[phab:T437563|T437563]]
* 05:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet,db1245.eqiad.wmnet with reason: Cloning
* 05:25 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] (duration: 09m 59s)
* 05:21 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:20 tstarling@deploy1003: tstarling: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:15 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 50s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-10 ==
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1349.eqiad.wmnet
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1349.eqiad.wmnet
* 23:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1349.eqiad.wmnet
* 23:09 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1349
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1349
* 22:32 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1349
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:28 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1349
* 22:27 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:27 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1349.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1348.eqiad.wmnet
* 22:23 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 22:12 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] (duration: 10m 59s)
* 22:06 musikanimal@deploy1003: kemayo, musikanimal: Rolling back deployment
* 22:05 musikanimal@deploy1003: kemayo, musikanimal: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:01 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]]
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:00 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:52 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:47 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:47 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] (duration: 13m 23s)
* 21:46 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:46 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2006.codfw.wmnet
* 21:42 derenrich@deploy1003: derenrich: Continuing with deployment
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1348
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1348
* 21:39 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1348
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:39 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:37 derenrich@deploy1003: derenrich: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:35 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:34 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1348
* 21:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 21:33 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1348.eqiad.wmnet
* 21:33 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]]
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1348.eqiad.wmnet
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1348.eqiad.wmnet
* 21:31 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2006.codfw.wmnet
* 21:31 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] (duration: 09m 45s)
* 21:27 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:26 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]]
* 21:19 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:19 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2006.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 21:17 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] (duration: 13m 54s)
* 21:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2005.codfw.wmnet with OS bookworm
* 21:12 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:07 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdeb
* 21:03 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]]
* 20:54 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:52 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] (duration: 23m 49s)
* 20:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:47 jdrewniak@deploy1003: jdrewniak, milazg: Continuing with deployment
* 20:34 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1346.eqiad.wmnet
* 20:32 jdrewniak@deploy1003: jdrewniak, milazg: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2005.codfw.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:28 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]]
* 20:27 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2005.codfw.wmnet
* 20:26 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:24 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:22 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] (duration: 11m 24s)
* 20:17 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 20:15 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2005.codfw.wmnet
* 20:14 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:12 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 20:11 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:07 jdrewniak@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.18,1.47.0-wmf.19,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted
* 20:05 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:02 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:01 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2005.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 19:53 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:50 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1346
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 19:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1346.eqiad.wmnet
* 19:21 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 thcipriani: Gerrit downtime incoming for upgrade
* 19:17 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hadoop.roll-restart-workers (exit_code=0) restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 dzahn@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:30:00 on gerrit.wikimedia.org with reason: maintenance upgrade
* 19:16 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on gerrit2003.wikimedia.org with reason: maintenance upgrade
* 19:04 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 18:21 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e] (duration: 00m 59s)
* 18:20 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e]
* 18:19 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e] (duration: 05m 13s)
* 18:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 18:14 otto@deploy1003: Started deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e]
* 18:13 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e] (duration: 00m 39s)
* 18:13 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e]
* 18:13 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 18:12 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 18:11 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 18:11 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit2002.wikimedia.org with reason: maintenance upgrade
* 18:11 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:11 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 18:10 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit1003.wikimedia.org with reason: maintenance upgrade
* 18:09 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 18:08 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 18:06 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 16:40 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 16:35 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 16:33 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]] synced to the te
* 16:28 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 15:33 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4
* 15:04 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] (duration: 08m 08s)
* 15:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host clouddumps1001.wikimedia.org with OS bookworm
* 14:59 samtar@deploy1003: samtar: Continuing with deployment
* 14:58 samtar@deploy1003: samtar: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:56 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]]
* 14:40 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:07 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 13:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host clouddumps1001.wikimedia.org with OS bookworm
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:55 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:51 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:42 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:41 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:48 klausman@dns1004: END - running authdns-update
* 12:46 klausman@dns1004: START - running authdns-update
* 12:35 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 12:35 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 12:05 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning x4
* 12:01 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2005.codfw.wmnet
* 11:55 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2005.codfw.wmnet
* 11:54 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1144.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:52 cgoubert@dns1004: END - running authdns-update
* 11:49 cgoubert@dns1004: START - running authdns-update
* 11:31 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2004.codfw.wmnet
* 11:25 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2004.codfw.wmnet
* 11:24 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:16 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1200.eqiad.wmnet
* 11:16 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1200.eqiad.wmnet
* 11:04 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1200.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:04 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1199.eqiad.wmnet
* 11:03 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1199.eqiad.wmnet
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:42 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1199.eqiad.wmnet with reason: Upgrading RAID firmware
* 10:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1024.eqiad.wmnet with reason: Cloning x4
* 10:00 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 09:56 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 09:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1024.eqiad.wmnet
* 09:45 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 09:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 09:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:04 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:51 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:46 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning x4
* 08:43 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:34 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] (duration: 09m 56s)
* 08:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:30 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:29 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:28 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:24 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]]
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 08:09 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 07:58 XioNoX: netflow1004:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:56 XioNoX: netflow2005:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 07:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 07:03 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:59 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:43 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:42 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:39 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 06:38 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 06:37 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 05:04 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 05:03 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tool-server' for release 'main' .
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 38s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1345.eqiad.wmnet
* 00:11 dzahn@dns1004: END - running authdns-update
* 00:08 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 00:08 dzahn@dns1004: START - running authdns-update
== 2026-09-09 ==
* 23:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:29 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1345
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 23:28 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1345.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1344.eqiad.wmnet
* 23:22 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] (duration: 11m 15s)
* 23:18 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 23:16 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 23:11 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]]
* 22:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:51 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2004.codfw.wmnet with OS bookworm
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1344
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1344
* 22:37 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1344
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1344
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 22:32 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] (duration: 10m 21s)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1344.eqiad.wmnet
* 22:27 derenrich@deploy1003: derenrich, egardner: Continuing with deployment
* 22:26 derenrich@deploy1003: derenrich, egardner: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:24 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:22 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]]
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:19 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] (duration: 13m 40s)
* 22:16 derenrich@deploy1003: derenrich: Rolling back deployment
* 22:10 derenrich@deploy1003: derenrich: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:05 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]]
* 22:03 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:44 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:40 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:36 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1343
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1343
* 21:27 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1343
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:27 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:23 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:22 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1343
* 21:22 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] (duration: 12m 29s)
* 21:21 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 21:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1343.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1342.eqiad.wmnet
* 21:17 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:14 jforrester@deploy1003: jforrester: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:09 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]]
* 21:08 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:07 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 21:07 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:06 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2004.codfw.wmnet
* 20:59 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:57 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:56 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2004.codfw.wmnet
* 20:53 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:53 bking@cumin2003: END (ERROR) - Cookbook sre.presto.roll-restart-workers (exit_code=97) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:53 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:50 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* {{safesubst:SAL entry|1=20:45 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2}}
* 20:42 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* 20:41 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 20:40 sbassett@deploy1003: aranyap, sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:39 sbassett@deploy1003: aranyap, sbassett: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "Filter}}
* 20:35 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* {{safesubst:SAL entry|1=20:34 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "}}
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:28 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:24 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 20:24 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1342
* 20:23 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 20:23 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1342.eqiad.wmnet
* 20:23 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1342.eqiad.wmnet
* 20:22 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1342.eqiad.wmnet
* 20:15 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1027.eqiad.wmnet with OS bookworm
* 19:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:41 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1027.eqiad.wmnet with OS bookworm
* 19:36 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:28 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:26 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:19 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 19:19 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:06 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 19:05 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 19:03 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 18:59 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 18:23 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:06 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1026.eqiad.wmnet with OS bookworm
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 18:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 17:54 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 17:49 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 17:45 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1026.eqiad.wmnet with OS bookworm
* 17:31 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:29 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 17:27 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 17:12 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 17:04 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 16:46 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] (duration: 09m 28s)
* 16:41 urbanecm@deploy1003: migr, urbanecm: Continuing with deployment
* 16:41 urbanecm@deploy1003: migr, urbanecm: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:36 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]]
* 16:28 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:25 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 15:55 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 15:54 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 15:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] (duration: 09m 43s)
* 15:41 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 15:40 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:36 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]]
* 15:17 urbanecm: Delete all running periodic jobs starting with `growthexperiments-refreshlinkrecommendations-*` (to pick up new configuration; [[phab:T392944|T392944]])
* 15:08 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:06 moritzm: installing grub2 bugfix updates from Bookworm point release
* 15:04 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp6008.drmrs.wmnet
* 15:01 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:42 hnowlan: half concurrency for parsoidCachePrewarm RecordLintJob and refreshLinks in jobqueue, eqiad & codfw
* 14:35 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 14:34 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:32 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 14:20 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:07 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2013.codfw.wmnet with OS trixie
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 14:06 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 13:55 btullis@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'.
* 13:53 btullis@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'.
* 13:43 btullis@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 13:42 btullis@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 13:29 moritzm: pruned obsolete Bullseye image dispatch from the docker registry [[phab:T416452|T416452]]
* 13:28 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:26 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b7-eqiad
* 13:25 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 13:24 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:22 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 13:22 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:17 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a4-eqiad
* 13:17 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] (duration: 10m 15s)
* 13:14 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 13:11 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 13:10 sbisson@deploy1003: sbisson: Continuing with deployment
* 13:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:04 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]]
* 13:02 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on ldap-rw[1001,2001].wikimedia.org with reason: work in progress
* 12:49 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 12:48 btullis@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
* 12:46 btullis@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
* 12:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] (duration: 14m 39s)
* 12:41 ladsgroup@deploy1003: tryvix1509, ladsgroup: Continuing with deployment
* 12:35 ladsgroup@deploy1003: tryvix1509, ladsgroup: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:31 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]]
* 12:16 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 12:16 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 11:53 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] (duration: 21m 58s)
* 11:48 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 11:35 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:31 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]]
* 10:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Repooling db2196
* 10:47 moritzm: pruned obsolete Bullseye images nodejs12-slim/nodejs12-devel/nodejs14-slim/nodejs16-slim from the docker registry [[phab:T416452|T416452]]
* 10:43 moritzm: installing Bird security updates
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1260: Repooling after cloning
* 10:09 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Repooling db2196
* 10:07 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 10:06 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 09:55 moritzm: pruned obsolete Bullseye images openjdk-8-jdk/openjdk-8-jre/openjdk-11-jre/openjdk-11-jdk from the docker registry [[phab:T416452|T416452]]
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1260: Repooling after cloning
* 09:52 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 09:52 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 09:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:28 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:27 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:03 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 09:02 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 09:01 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 16 hosts with reason: upgrade ssw1-a1-eqiad
* 08:58 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 22 hosts with reason: upgrade ssw1-a1-eqiad
* 08:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 08:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 08:49 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 08:49 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 08:48 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 08:48 cmooney@cumin1004: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 08:36 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning sanitarium
* 08:30 brouberol@dns1004: END - running authdns-update
* 08:28 moritzm: pruned obsolete Bullseye image golang1.15 from the docker registry [[phab:T416452|T416452]]
* 08:28 brouberol@dns1004: START - running authdns-update
* 08:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1260: Needs to clone another host from this one
* 08:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1260: Needs to clone another host from this one
* 08:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1260.eqiad.wmnet with reason: Cloning sanitarium
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:00 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 07:40 chlod: UTC morning backport window done
* 07:37 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] (duration: 21m 36s)
* 07:32 chlod@deploy1003: chlod, hamishz: Continuing with deployment
* 07:20 chlod@deploy1003: chlod, hamishz: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 45s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1025.eqiad.wmnet with OS bookworm
== 2026-09-08 ==
* 23:51 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:48 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:39 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1313.eqiad.wmnet
* 23:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:25 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 23:19 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:19 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:15 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:05 Amir1: dropped 57 tables on db1260 ([[phab:T437278|T437278]])
* 23:03 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 23:03 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:02 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 22:57 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1313
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1313
* 22:37 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1313
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:37 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:33 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 22:33 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1313
* 22:32 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 22:32 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1313.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1306.eqiad.wmnet
* 22:27 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6008.drmrs.wmnet with OS trixie
* 22:18 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 22:03 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 22:01 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] (duration: 09m 53s)
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 21:58 Amir1: drop links tables from db2210 ([[phab:T437278|T437278]])
* 21:57 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 21:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:56 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:52 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 21:51 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]]
* 21:51 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:48 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:45 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] (duration: 05m 27s)
* 21:43 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Continuing with deployment
* 21:40 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:39 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]]
* 21:35 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1024.eqiad.wmnet with OS bookworm
* 21:33 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp6008.drmrs.wmnet with OS trixie
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1306
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1306
* 21:30 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1306
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:24 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] (duration: 09m 12s)
* 21:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 21:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1306
* 21:19 reedy@deploy1003: reedy: Continuing with deployment
* 21:19 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:19 reedy@deploy1003: reedy: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1306.eqiad.wmnet
* 21:15 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:15 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]]
* {{safesubst:SAL entry|1=21:10 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out}}
* 21:05 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1024.eqiad.wmnet with OS bookworm
* 21:05 sbassett@deploy1003: sbassett: Continuing with deployment
* 21:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 21:04 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=21:03 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out non-http(s) lice}}
* 20:59 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=20:58 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out n}}
* 20:53 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 20:50 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1305.eqiad.wmnet
* 20:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 20:34 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 20:28 sbassett@deploy1003: sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:27 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-http(s) license}}
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1023.eqiad.wmnet with OS bookworm
* {{safesubst:SAL entry|1=20:23 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-}}
* 20:15 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] (duration: 10m 16s)
* 20:13 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:10 aaron@deploy1003: aaron: Continuing with deployment
* 20:09 aaron@deploy1003: aaron: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 20:05 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:05 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]]
* 20:01 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1023.eqiad.wmnet with OS bookworm
* 19:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1305
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1305
* 19:43 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1305
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:43 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:39 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 19:39 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1305
* 19:38 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 19:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1305.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1275.eqiad.wmnet
* 19:23 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 19:17 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 18:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:53 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1275
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1275
* 18:33 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1275
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1275
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 18:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1275.eqiad.wmnet
* 18:18 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 17:43 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy1003.eqiad.wmnet
* 17:36 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet
* 17:36 cdobbins@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-ntp (exit_code=0) rolling restart_daemons on A:dnsbox
* 17:30 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy1003.eqiad.wmnet
* 17:25 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet
* 17:15 swfrench@deploy1003: Finished scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup (duration: 04m 18s)
* 17:11 Amir1: dropping links tables from db1247 (s4 replica) - ([[phab:T437278|T437278]])
* 17:10 swfrench@deploy1003: Started scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup
* 16:51 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] (duration: 10m 19s)
* 16:46 zabe@deploy1003: zabe: Continuing with deployment
* 16:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]]
* 16:29 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 jhancock@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:12 jhancock@cumin2003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 16:08 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2003.codfw.wmnet
* 15:58 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:55 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2003.codfw.wmnet
* 15:54 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:44 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2002.codfw.wmnet
* 15:19 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2002.codfw.wmnet
* 14:55 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:46 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2001.codfw.wmnet
* 14:44 topranks: shutdown et-1/1/5 on cr1-codfw to shift traffic off ssw1-a1-codfw
* 14:43 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: upgrade ssw1-a1-eqiad
* 14:34 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2001.codfw.wmnet
* 14:33 btullis@cumin1003: END (ERROR) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=97) rolling reboot on A:cephosd-codfw
* 14:30 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet
* 14:28 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet
* 14:22 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet
* 14:13 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --olderThan {{Gerrit|1788220800}} --verbose # [[phab:T437158|T437158]]
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet
* 14:12 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet
* 14:03 topranks: drain traffic from ssw1-a1-codfw before JunOS upgrade [[phab:T426197|T426197]]
* 14:02 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet
* 13:58 cgoubert@deploy1003: helmfile [staging-codfw] DONE helmfile.d/services/mw-debug: apply
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet
* 13:57 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet
* 13:56 cgoubert@deploy1003: helmfile [staging-codfw] START helmfile.d/services/mw-debug: apply
* 13:55 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:54 stran@deploy1003: mwscript-k8s job started: foreachwikiindblist checkuser-suggested-investigations extensions/CheckUser/maintenance/populateSiCaseProperties.php # [[phab:T435066|T435066]]
* 13:52 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:51 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet
* 13:50 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:46 cdobbins@cumin1003: START - Cookbook sre.dns.roll-restart-ntp rolling restart_daemons on A:dnsbox
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet
* 13:46 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet
* 13:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2012.codfw.wmnet with OS bookworm
* 13:44 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] (duration: 34m 00s)
* 13:40 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet
* 13:37 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet
* 13:35 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet
* 13:32 stran@deploy1003: stran: Continuing with deployment
* 13:29 stran@deploy1003: stran: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet
* 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:21 moritzm: installing qemu security updates
* 13:18 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet
* 13:16 ayounsi@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'email' for AS: 139628
* 13:15 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 139628
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet
* 13:13 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet
* 13:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 2519
* 13:11 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 2519
* 13:10 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 14593
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]]
* 13:09 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:08 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 14593
* 13:06 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet
* 13:06 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2012.codfw.wmnet with OS bookworm
* 13:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2012.codfw.wmnet
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2012.codfw.wmnet
* 13:05 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:04 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet
* 13:01 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet
* 12:58 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 34655
* 12:58 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 34655
* 12:56 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2012.codfw.wmnet
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'clear' for AS: 35320
* 12:55 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'clear' for AS: 35320
* 12:55 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-codfw
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-codfw
* 12:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2011.codfw.wmnet with OS bookworm
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-by27-esams
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-by27-esams
* 12:54 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-esams
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-esams
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-esams
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-esams
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-eqsin
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-eqsin
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-ulsfo
* 12:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f3-eqiad
* 12:50 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f3-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-c8-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-c8-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:44 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:43 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet
* 12:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-eqiad
* 12:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f1-eqiad
* 12:41 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f1-eqiad
* 12:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-d5-eqiad
* 12:40 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-d5-eqiad
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet
* 12:38 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet
* 12:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:32 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet
* 12:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad
* 12:11 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2012.codfw.wmnet
* 12:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2011.codfw.wmnet with OS bookworm
* 12:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2011.codfw.wmnet
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2011.codfw.wmnet
* 12:09 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:06 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] (duration: 09m 54s)
* 12:01 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2011.codfw.wmnet
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Continuing with deployment
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2010.codfw.wmnet with OS bookworm
* 11:56 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]]
* 11:46 marostegui@dns1004: END - running authdns-update
* 11:44 marostegui@dns1004: START - running authdns-update
* 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:40 Amir1: dropping unneeded tables from x4 - db1260 ([[phab:T437278|T437278]])
* 11:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:23 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2011.codfw.wmnet
* 11:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2010.codfw.wmnet with OS bookworm
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2010.codfw.wmnet
* 11:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2010.codfw.wmnet
* 11:20 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2010.codfw.wmnet
* 11:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2009.codfw.wmnet with OS bookworm
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:43 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] (duration: 10m 57s)
* 10:39 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:38 samtar@deploy1003: samtar: Continuing with deployment
* 10:37 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:37 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:36 samtar@deploy1003: samtar: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:34 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:32 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]]
* 10:30 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2009.codfw.wmnet with OS bookworm
* 10:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2009.codfw.wmnet
* 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:18 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2009.codfw.wmnet
* 10:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2008.codfw.wmnet with OS bookworm
* 10:07 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:05 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:01 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:45 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] (duration: 13m 15s)
* 09:45 ayounsi@dns1004: END - running authdns-update
* 09:44 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:43 ayounsi@dns1004: START - running authdns-update
* 09:39 zabe@deploy1003: zabe: Continuing with deployment
* 09:37 zabe@deploy1003: zabe: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2008.codfw.wmnet with OS bookworm
* 09:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:32 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]]
* 09:32 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2008.codfw.wmnet
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2008.codfw.wmnet
* 09:31 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:29 XioNoX: remove GRE tunnels eqiad-drmrs eqdfw-ulsfo
* 09:23 moritzm: installing rsync security updates
* 09:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2007.codfw.wmnet with OS bookworm
* 09:22 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2008.codfw.wmnet
* 09:11 marostegui@cumin1003: dbctl commit (dc=all): 'Make x4 and s4 RW again [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96393 and previous config saved to /var/cache/conftool/dbconfig/20260908-091121-marostegui.json
* 09:07 marostegui@cumin1003: dbctl commit (dc=all): 'Remove old s4 masters from x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96392 and previous config saved to /var/cache/conftool/dbconfig/20260908-090749-marostegui.json
* 09:05 marostegui@cumin1003: dbctl commit (dc=all): 'Set x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96391 and previous config saved to /var/cache/conftool/dbconfig/20260908-090517-marostegui.json
* 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Set s4 commons to read-only for maintenance [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96389 and previous config saved to /var/cache/conftool/dbconfig/20260908-090228-marostegui.json
* 09:02 marostegui: Starting x4 split from s4, RO time on commons needed [[phab:T404715|T404715]]
* 09:00 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 08:58 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2008.codfw.wmnet
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:43 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 32 hosts with reason: x4 split
* 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2007.codfw.wmnet with OS bookworm
* 08:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:37 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:37 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2007.codfw.wmnet
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2007.codfw.wmnet
* 08:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:34 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:29 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:26 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2007.codfw.wmnet
* 08:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2006.codfw.wmnet with OS bookworm
* 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:50 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2007.codfw.wmnet
* 07:43 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2006.codfw.wmnet with OS bookworm
* 07:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2006.codfw.wmnet
* 07:37 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:30 denisse: Add grafana-plugins 0.15 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 07:29 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:27 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 07:27 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:27 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:22 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 07:22 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 07:18 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:14 jmm@dns1004: END - running authdns-update
* 07:13 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96388 and previous config saved to /var/cache/conftool/dbconfig/20260908-071308-marostegui.json
* 07:12 jmm@dns1004: START - running authdns-update
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96387 and previous config saved to /var/cache/conftool/dbconfig/20260908-071216-marostegui.json
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96386 and previous config saved to /var/cache/conftool/dbconfig/20260908-071159-marostegui.json
* 05:07 denisse: Add grafana-plugins 0.10 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.16 (duration: 02m 27s)
* 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]] (duration: 36m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 41s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-07 ==
* 21:52 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] (duration: 11m 00s)
* 21:47 zabe@deploy1003: zabe: Continuing with deployment
* 21:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]]
* 21:37 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] (duration: 09m 34s)
* 21:33 zabe@deploy1003: zabe: Continuing with deployment
* 21:32 zabe@deploy1003: zabe: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:28 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]]
* 21:03 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] (duration: 10m 27s)
* 20:58 zabe@deploy1003: zabe: Continuing with deployment
* 20:57 zabe@deploy1003: zabe: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:52 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]]
* 20:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set weight of db1261 to zero in s4 ([[phab:T437108|T437108]])', diff saved to https://phabricator.wikimedia.org/P96385 and previous config saved to /var/cache/conftool/dbconfig/20260907-203804-ladsgroup.json
* 20:23 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] (duration: 09m 28s)
* 20:19 zabe@deploy1003: zabe: Continuing with deployment
* 20:18 zabe@deploy1003: zabe: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:14 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]]
* 20:12 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] (duration: 10m 06s)
* 20:07 zabe@deploy1003: zabe, daimona: Continuing with deployment
* 20:06 zabe@deploy1003: zabe, daimona: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:02 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]]
* 19:59 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] (duration: 11m 27s)
* 19:55 zabe@deploy1003: zabe: Continuing with deployment
* 19:52 zabe@deploy1003: zabe: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:48 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]]
* 19:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] (duration: 11m 40s)
* 19:26 zabe@deploy1003: zabe: Continuing with deployment
* 19:23 zabe@deploy1003: zabe: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:19 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]]
* 19:07 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] (duration: 14m 17s)
* 19:00 zabe@deploy1003: zabe: Continuing with deployment
* 18:57 zabe@deploy1003: zabe: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:53 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]]
* 18:33 zabe@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 18:32 zabe@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 18:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] (duration: 09m 12s)
* 18:27 zabe@deploy1003: zabe: Continuing with deployment
* 18:26 zabe@deploy1003: zabe: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:22 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]]
* 16:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2005.codfw.wmnet with OS bookworm
* 16:01 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] (duration: 10m 22s)
* 15:59 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 15:57 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 15:56 zabe@deploy1003: zabe: Continuing with deployment
* 15:55 zabe@deploy1003: zabe: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:51 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]]
* 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:47 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 15:46 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 15:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:44 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 15:44 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 15:27 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2005.codfw.wmnet with OS bookworm
* 15:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2005.codfw.wmnet
* 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:11 moritzm: installing rsync security updates
* 15:04 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1341.eqiad.wmnet
* 15:03 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2004.codfw.wmnet with OS bookworm
* 15:00 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:58 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 14:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:40 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 14:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 14:35 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:32 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:29 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:25 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 14:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1228: Repooling db1228 into s4
* 14:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2004.codfw.wmnet with OS bookworm
* 14:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Repooling after cloning
* 14:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:19 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2004.codfw.wmnet
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2004.codfw.wmnet
* 14:18 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:14 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1074.eqiad.wmnet
* 14:14 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:13 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:13 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* {{safesubst:SAL entry|1=14:11 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mul}}
* 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2004.codfw.wmnet
* 14:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1074.eqiad.wmnet
* 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1073.eqiad.wmnet
* 14:07 krinkle@deploy1003: krinkle: Continuing with deployment
* {{safesubst:SAL entry|1=14:04 krinkle@deploy1003: krinkle: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with multiple properties}}
* 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1073.eqiad.wmnet
* 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1072.eqiad.wmnet
* 14:01 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1341
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1341
* 14:01 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 14:00 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1341
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* 13:59 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* {{safesubst:SAL entry|1=13:59 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mult}}
* 13:59 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2004.codfw.wmnet
* 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2003.codfw.wmnet with OS bookworm
* 13:55 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 13:55 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1072.eqiad.wmnet
* 13:55 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudvirt1067.eqiad.wmnet
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1341
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 13:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1341.eqiad.wmnet
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2008.wikimedia.org
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2008.wikimedia.org with OS trixie
* 13:39 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1067.eqiad.wmnet
* 13:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1066.eqiad.wmnet
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:37 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1228: Repooling db1228 into s4
* 13:36 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Repooling after cloning
* 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1066.eqiad.wmnet
* 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1065.eqiad.wmnet
* 13:28 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:27 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1065.eqiad.wmnet
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:25 moritzm: installing openssh security updates
* 13:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:24 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1340.eqiad.wmnet
* 13:23 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] (duration: 10m 06s)
* 13:17 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2003.codfw.wmnet with OS bookworm
* 13:17 stran@deploy1003: stran: Continuing with deployment
* 13:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:16 stran@deploy1003: stran: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:16 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:15 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 13:12 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]]
* 13:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2003.codfw.wmnet
* 13:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2003.codfw.wmnet
* 13:03 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2008.wikimedia.org with OS trixie
* 13:03 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:02 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:02 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:01 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:00 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2003.codfw.wmnet
* 12:58 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:54 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 12:54 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2008.wikimedia.org
* 12:47 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2003.codfw.wmnet
* 12:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2002.codfw.wmnet with OS bookworm
* 12:29 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 12:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2007.wikimedia.org
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2007.wikimedia.org with OS trixie
* 12:14 elukey: moved most of the Docker Registry's prefixes to a new internal S3 backend. For any docker pull failure that worked in the past, please ping me or drop a note in [[phab:T435499|T435499]] or contact the oncall SREs
* 12:07 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2002.codfw.wmnet with OS bookworm
* 12:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:02 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2002.codfw.wmnet
* 11:54 jmm@dns1004: END - running authdns-update
* 11:52 jmm@dns1004: START - running authdns-update
* 11:46 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2002.codfw.wmnet
* 11:46 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2007.wikimedia.org with OS trixie
* 11:46 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:46 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:41 moritzm: installing bash updates from bookworm point release
* 11:39 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:39 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2007.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts ldap-replica1006.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:35 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:32 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2001.codfw.wmnet with OS bookworm
* 11:28 jmm@cumin1004: START - Cookbook sre.hosts.decommission for hosts ldap-replica1006.wikimedia.org
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 11:18 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] (duration: 14m 08s)
* 11:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 11:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet
* 11:11 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:11 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:10 zabe@deploy1003: zabe: Continuing with deployment
* 11:10 zabe@deploy1003: zabe: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 11:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet
* 11:07 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet
* 11:07 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet
* 11:06 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:06 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:04 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]]
* 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 10:58 btullis@deploy1003: Finished scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]] (duration: 35m 20s)
* 10:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:53 jmm@dns1004: END - running authdns-update
* 10:51 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:50 jmm@dns1004: START - running authdns-update
* 10:47 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=frwiki # [[phab:T436659|T436659]]
* 10:40 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=hrwiki # [[phab:T436659|T436659]]
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:39 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 10:37 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 10:33 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 10:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2001.codfw.wmnet with OS bookworm
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1340
* 10:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1340.eqiad.wmnet
* 10:26 btullis@deploy1003: Started scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]]
* 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2001.codfw.wmnet
* 10:18 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2001.codfw.wmnet
* 10:13 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw
* 10:12 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:10 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:09 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:08 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:07 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:07 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:02 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:00 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:59 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:57 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/services/miscweb: apply
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/services/miscweb: apply
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:52 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:52 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] (duration: 10m 11s)
* 09:51 aokoth@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 09:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 09:49 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 09:49 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 09:48 aokoth@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 09:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:45 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:44 aokoth@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 09:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]]
* 09:41 aokoth@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 09:38 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Repooling after switchover
* 09:23 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw
* 09:15 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] (duration: 20m 12s)
* 09:12 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 139009
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1005.wikimedia.org
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1005.wikimedia.org with OS trixie
* 09:10 moritzm: rebuild software RAID following disk replacement [[phab:T437036|T437036]]
* 09:10 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 139009
* 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1022.eqiad.wmnet with OS bookworm
* 09:08 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 09:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:03 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet
* 09:02 moritzm: installing giflib security updates
* 09:01 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:59 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 08:57 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 08:56 jmm@cumin1004: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet
* 08:56 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:54 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]]
* 08:52 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:52 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:49 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 08:48 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:42 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:41 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:40 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:40 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96355 and previous config saved to /var/cache/conftool/dbconfig/20260907-083904-marostegui.json
* 08:38 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1005.wikimedia.org with OS trixie
* 08:38 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2157 to s5 primary [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96354 and previous config saved to /var/cache/conftool/dbconfig/20260907-083825-marostegui.json
* 08:38 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 marostegui: Starting s5 codfw failover from db2213 to db2157 - [[phab:T437188|T437188]]
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2157 with weight 0 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96353 and previous config saved to /var/cache/conftool/dbconfig/20260907-083448-marostegui.json
* 08:34 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s5 [[phab:T437188|T437188]]
* 08:28 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:28 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1005.wikimedia.org
* 08:22 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:20 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1022.eqiad.wmnet with OS bookworm
* 08:03 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:02 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 08:02 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:00 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 07:57 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 07:56 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1241.eqiad.wmnet with reason: Cloning
* 07:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1241: Cloning
* 07:55 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1241: Cloning
* 07:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Cloning
* 07:54 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Cloning
* 07:51 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 07:50 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 07:47 kartik@deploy1003: Finished scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] (duration: 41m 51s)
* 07:46 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 07:45 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 07:34 kartik@deploy1003: abi, kartik: Continuing with deployment
* 07:23 kartik@deploy1003: abi, kartik: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:05 kartik@deploy1003: Started scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]]
* 06:14 moritzm: installing Chromium security updates
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 08m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-06 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 25s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-05 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 26s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-04 ==
* 22:07 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 21:48 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:42 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:30 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:42 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:40 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 19:27 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:19 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:13 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:12 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host sretest2013
* 19:11 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host sretest2013
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:11 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:07 jhancock@cumin1003: START - Cookbook sre.dns.netbox
* 18:18 inflatador: bking@clouddumps100[12] `systemctl reset-failed` to quash alerts until https://w.wiki/UBje . The systemd timer should try again tomorrow
* 17:27 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b8-eqiad
* 17:27 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b8-eqiad
* 16:37 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:33 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-b7-eqiad
* 16:33 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 16:05 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b6-eqiad
* 16:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b6-eqiad
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1339.eqiad.wmnet
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 15:50 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:49 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:47 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:43 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 15:10 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b5-eqiad
* 15:09 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b5-eqiad
* 14:46 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1045.eqiad.wmnet
* 14:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1003.eqiad.wmnet with OS trixie
* 14:38 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 14:37 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b4-eqiad
* 14:37 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b4-eqiad
* 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339
* 14:32 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339
* 14:31 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 14:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet
* 14:26 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:24 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:21 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b3-eqiad
* 14:21 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b3-eqiad
* 14:17 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:17 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:04 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1003.eqiad.wmnet with OS trixie
* 13:54 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b2-eqiad
* 13:53 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b2-eqiad
* 13:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1002.eqiad.wmnet with OS trixie
* 13:18 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a4-eqiad
* 13:12 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:12 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b1-eqiad
* 13:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b1-eqiad
* 12:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1002.eqiad.wmnet with OS trixie
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow3004.esams.wmnet with OS trixie
* 12:33 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:28 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 12:15 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 12:11 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a7-eqiad
* 12:11 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a7-eqiad
* 12:01 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow3004.esams.wmnet with OS trixie
* 11:47 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a6-eqiad
* 11:47 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a6-eqiad
* 11:36 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki2003.codfw.wmnet
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki2003.codfw.wmnet with OS trixie
* 11:19 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:13 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a5-eqiad
* 11:06 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a5-eqiad
* 10:52 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki2003.codfw.wmnet with OS trixie
* 10:50 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:50 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:44 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 10:44 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki2003.codfw.wmnet
* 10:29 btullis@deploy1003: Finished scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli (duration: 41m 14s)
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 10:00 marostegui@cumin1003: Removing db1182 from zarcillo [[phab:T434869|T434869]]
* 10:00 marostegui@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db1182.eqiad.wmnet
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:57 btullis@deploy1003: Started scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli
* 09:53 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:53 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.decommission (exit_code=1)
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1182.eqiad.wmnet
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:50 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:46 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:45 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1182 from dbctl [[phab:T434869|T434869]]', diff saved to https://phabricator.wikimedia.org/P96346 and previous config saved to /var/cache/conftool/dbconfig/20260904-094527-marostegui.json
* 09:41 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki1003.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki1003.eqiad.wmnet with OS trixie
* 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1182: Decommissioning
* 09:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1182: Decommissioning
* 09:23 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 09:17 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow5003.eqsin.wmnet with OS trixie
* 09:02 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki1003.eqiad.wmnet with OS trixie
* 09:00 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 09:00 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:55 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:55 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki1003.eqiad.wmnet
* 08:54 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:48 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:45 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:40 btullis@deploy1003: Finished scap sync-world: Trying again for [[phab:T436913|T436913]] (duration: 34m 26s)
* 08:35 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2003.codfw.wmnet
* 08:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw2001.wikimedia.org with OS trixie
* 08:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4 days, 0:00:00 on db2196.codfw.wmnet with reason: Host crashed
* 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2003.codfw.wmnet
* 08:24 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:20 elukey@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:12 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:08 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:07 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 08:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:04 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 07:57 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow5003.eqsin.wmnet with OS trixie
* 07:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:55 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw2001.wikimedia.org with OS trixie
* 07:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:51 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:50 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie
* 07:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2196: down
* 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2196: down
* 06:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:52 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:40 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:28 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:21 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 06:20 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 39s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 01:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:02 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:59 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:46 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 00:44 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:39 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:23 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
== 2026-09-03 ==
* 21:46 tsev@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T435363|T435363]]
* 21:03 eevans@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:55 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] (duration: 12m 24s)
* 20:52 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:52 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1021.eqiad.wmnet with OS bookworm
* 20:50 arlolra@deploy1003: arlolra, tsev: Continuing with deployment
* 20:46 arlolra@deploy1003: arlolra, tsev: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:44 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1047.eqiad.wmnet
* 20:42 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]]
* 20:41 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a3-eqiad
* 20:40 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a3-eqiad
* 20:39 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] (duration: 10m 23s)
* 20:34 arlolra@deploy1003: arlolra, jgiannelos: Continuing with deployment
* 20:33 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1047.eqiad.wmnet
* 20:33 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:32 arlolra@deploy1003: arlolra, jgiannelos: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:31 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1046.eqiad.wmnet
* 20:28 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]]
* 20:23 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] (duration: 13m 41s)
* 20:20 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1046.eqiad.wmnet
* 20:16 catrope@deploy1003: catrope: Continuing with deployment
* 20:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1021.eqiad.wmnet with OS bookworm
* 20:13 catrope@deploy1003: catrope: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1021.eqiad.wmnet
* 20:09 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]]
* 20:00 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1021.eqiad.wmnet
* 19:49 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 19:19 swfrench@deploy1003: Finished scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]] (duration: 02m 59s)
* 19:16 swfrench@deploy1003: Started scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]]
* 19:01 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:01 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 18:57 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:57 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:34 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:20 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 17:55 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:55 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:53 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:52 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:46 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 17:46 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 17:45 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 17:44 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 17:44 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:40 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:39 ryankemper: [WDQS] Service looks healthy again, CPU load and thread count have dropped considerably over the last hour
* 17:39 ryankemper: [[phab:T421642|T421642]] [WDQS] requestctl changes: `2026-09-03 16:23-17:33` UTC: added hard-deny pair `cache-text/wdqs_futile_sparql_sep_2026_deny(+_bots)`; extended pattern `ua/wdqs_heavy_sparql_bots_2026` and added default-scope twin `wdqs_heavy_sparql_bots_jul_2026_ratelimit_default`; added ipblock `abuse/wdqs_sparql_scanners_sep_2026` + throttle `wdqs_sparql_scanners_sep_2026_ratelimit` (needed manual `requestctl update-provenance-map`)
* 17:37 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:37 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:31 andrew@cumin2003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudcephosd1045.eqiad.wmnet
* 17:30 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:30 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:28 dancy@deploy1003: Installation of scap version "4.289.0" completed for 3 hosts
* 17:26 dancy@deploy1003: Installing scap version "4.289.0" for 3 host(s)
* 17:24 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a2-eqiad
* 17:24 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:22 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:22 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a2-eqiad
* 17:14 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 17:10 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:10 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:09 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 17:01 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:59 andrew@cumin2003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 16:58 dancy: Running scap clean-images on deploy1003
* 16:52 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:50 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:39 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:14 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs2021.codfw.wmnet,service=wdqs-main
* 16:14 ryankemper: [[phab:T430880|T430880]] Stumbled across `wdqs2021` listed as inactive, looks like it was never fully re-pooled after a data xfer. Pooled.
* 16:12 btullis@deploy1003: Finished deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2] (duration: 00m 38s)
* 16:12 btullis@deploy1003: Started deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2]
* 16:12 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:07 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs101[1-4].eqiad.wmnet
* 16:03 btullis@deploy1003: Started scap sync-world: Rebuilding to pick up new version of dump scripts in mediawiki-cli for [[phab:T436913|T436913]]
* 16:01 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]] (duration: 09m 29s)
* 15:51 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=urldownloader[12]00[56].wikimedia.org [reason: depooling urldownloader trixie nodes]
* 15:51 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]]
* 15:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:29 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:24 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:24 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:18 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:15 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:09 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:05 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 15:04 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:02 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 15:00 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader
* 14:58 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw
* 14:53 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-ulsfo
* 14:49 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 14:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2002.codfw.wmnet
* 14:45 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:39 sukhe: sudo cumin "A:cp-text" "run-puppet-agent --enable 'merging CR 1334855'": [[phab:T425441|T425441]]
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2002.codfw.wmnet
* 14:32 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:30 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:29 sukhe: sudo cumin "A:cp-text" "disable-puppet 'merging CR 1334855'" [[phab:T425441|T425441]]
* 14:27 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-ulsfo
* 14:22 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:21 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:19 arnaudb@dns1006: END - running authdns-update
* 14:18 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1074
* 14:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1074
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-codfw
* 14:17 arnaudb@dns1006: START - running authdns-update
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:16 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:16 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1073
* 14:15 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1073
* 14:12 ayounsi@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host netflow2004.codfw.wmnet with OS trixie
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:10 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:08 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:08 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:07 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] (duration: 09m 36s)
* 14:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:05 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:03 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 14:03 samtar@deploy1003: btullis, samtar: Continuing with deployment
* 14:02 samtar@deploy1003: btullis, samtar: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:01 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:58 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]]
* 13:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:56 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:55 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:52 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqiad and A:durum
* 13:52 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-codfw
* 13:51 moritzm: installing sqlite3 security updates
* 13:51 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:51 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqiad and A:durum
* 13:49 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-codfw and A:durum
* 13:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:47 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-codfw and A:durum
* 13:47 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-esams
* 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:44 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:43 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:43 ayounsi@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:42 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] (duration: 13m 50s)
* 13:41 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:40 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1072
* 13:40 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:39 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-esams and A:durum
* 13:38 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1072
* 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:38 samtar@deploy1003: samtar, thiemowmde: Continuing with deployment
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-esams and A:durum
* 13:37 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:33 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:33 samtar@deploy1003: samtar, thiemowmde: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:32 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqsin and A:durum
* 13:31 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqsin and A:durum
* 13:28 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]]
* 13:28 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 13:27 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:24 ayounsi@cumin1004: START - Cookbook sre.hosts.reimage for host netflow2004.codfw.wmnet with OS trixie
* 13:24 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:22 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 13:22 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-esams
* 13:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:15 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:15 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 moritzm: installing bash updates from trixie point release
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 13:14 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1067
* 13:13 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1067
* 13:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2901: Test
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:09 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:09 moritzm: installing libxslt bugfix updates from Trixie point release
* 13:08 jelto@dns1004: END - running authdns-update
* 13:06 jelto@dns1004: START - running authdns-update
* 13:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 13:04 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:04 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 13:00 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:59 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1066
* 12:59 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1066
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:58 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:54 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 12:53 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2901: Test
* 12:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2901: Test
* 12:52 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'sync'.
* 12:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 12:50 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2901: Test
* 12:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:50 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 12:42 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:41 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-magru and A:durum
* 12:38 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-magru and A:durum
* 12:34 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 12:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1282: Pooling db1282 into s6
* 12:31 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 12:25 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:21 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:19 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:15 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_ulsfo
* 12:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1020.eqiad.wmnet with OS bookworm
* 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: db2207 repool
* 12:07 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_ulsfo
* 12:04 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 11:58 kart_: cxserver: Use urldownloader LVS endpoint ([[phab:T429175|T429175]])
* 11:57 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply
* 11:56 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply
* 11:56 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_magru
* 11:55 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply
* 11:55 moritzm: installing rsync security updates
* 11:55 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply
* 11:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:52 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply
* 11:51 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply
* 11:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1282: Pooling db1282 into s6
* 11:45 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1282 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96328 and previous config saved to /var/cache/conftool/dbconfig/20260903-114526-marostegui.json
* 11:43 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqiad
* 11:35 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqiad
* 11:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1020.eqiad.wmnet with OS bookworm
* 11:26 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 11:24 cgoubert@deploy1003: Finished scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter (duration: 12m 01s)
* 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: db2207 repool
* 11:22 cgoubert@deploy1003: cgoubert: Continuing with deployment
* 11:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqsin
* 11:17 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_esams
* 11:15 cgoubert@deploy1003: cgoubert: mediawiki: enable forward of fatal metrics to statsd exporter synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:14 cgoubert@deploy1003: Started scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter
* 11:10 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_esams
* 11:09 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 11:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1019.eqiad.wmnet with OS bookworm
* 11:01 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_drmrs
* 10:59 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_codfw
* 10:52 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_codfw
* 10:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:41 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 10:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 10:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow2005.codfw.wmnet
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2005.codfw.wmnet with OS trixie
* 10:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:17 btullis@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync
* 10:17 btullis@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync
* 10:16 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1019.eqiad.wmnet with OS bookworm
* 10:15 btullis@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: sync
* 10:15 btullis@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: sync
* 10:12 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s8
* 10:11 marostegui: Move s8 sanitarium from db1167 to db1281 [[phab:T434778|T434778]]
* 10:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 10:03 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 09:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:55 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2003.codfw.wmnet with OS trixie
* 09:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:43 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow2005.codfw.wmnet with OS trixie
* 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:42 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1018.eqiad.wmnet with OS bookworm
* 09:41 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_ulsfo
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:39 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2003.codfw.wmnet with reason: host reimage
* 09:39 hnowlan: fixed currently oncall pane in klaxon
* 09:38 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:38 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434751|T434751]]
* 09:37 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:37 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow2005.codfw.wmnet
* 09:35 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434775|T434775]]
* 09:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s7
* 09:34 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:33 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_ulsfo
* 09:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqiad
* 09:30 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] (duration: 09m 30s)
* 09:27 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:25 zabe@deploy1003: zabe: Continuing with deployment
* 09:25 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqiad
* 09:25 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 09:25 zabe@deploy1003: zabe: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1174 from dbctl [[phab:T436904|T436904]]', diff saved to https://phabricator.wikimedia.org/P96323 and previous config saved to /var/cache/conftool/dbconfig/20260903-092448-marostegui.json
* 09:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:21 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]]
* 09:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqsin
* 09:17 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:15 topranks: put traffic on Lumen codfw<->eqiad link as it is stable [[phab:T435810|T435810]]
* 09:14 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_esams
* 09:09 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:06 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_esams
* 09:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 09:03 marostegui: Move s6 sanitarium from db1165 to db1279 [[phab:T434775|T434775]]
* 09:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1018.eqiad.wmnet with OS bookworm
* 08:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s6
* 08:57 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_drmrs
* 08:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_codfw
* 08:55 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw
* 08:49 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_codfw
* 08:49 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:45 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 08:45 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:42 marostegui: Move s5 sanitarium from db1161 to db1275 [[phab:T434776|T434776]]
* 08:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 24 hosts with reason: Changing sanitarium master in s5
* 08:38 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_magru
* 08:37 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:37 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server (duration: 01m 21s)
* 08:37 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:36 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server
* 08:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:33 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server (duration: 01m 28s)
* 08:32 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server
* 08:27 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 08:11 moritzm: uploaded wmf-laptop 1.0.7 to apt.wikimedia.org
* 08:03 marostegui: Move s2 sanitarium from db1156 to db1271 [[phab:T434287|T434287]]
* 07:59 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2001.codfw.wmnet
* 07:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 23 hosts with reason: Changing sanitarium master in s2
* 07:49 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2001.codfw.wmnet
* 07:39 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:29 chlod: UTC morning backport window done
* 07:27 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] (duration: 11m 54s)
* 07:22 chlod@deploy1003: chlod, tryvix1509: Continuing with deployment
* 07:22 XioNoX: push pfw policies - [[phab:T436729|T436729]]
* 07:20 chlod@deploy1003: chlod, tryvix1509: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]]
* 07:15 marostegui: Power off db1228 for maintenance
* 07:13 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1228.eqiad.wmnet with reason: Onsite maintenance
* 07:01 arnaudb@dns1006: END - running authdns-update
* 06:58 arnaudb@dns1006: START - running authdns-update
* 06:54 jmm@cumin2003: END (PASS) - Cookbook sre.wdqs.restart-nginx-envoy (exit_code=0) rolling restart_daemons on A:wcqs-public
* 06:52 jmm@cumin2003: START - Cookbook sre.wdqs.restart-nginx-envoy rolling restart_daemons on A:wcqs-public
* 06:46 moritzm: installing libxml2 security updates
* 06:27 hashar: Upgrading CI Jenkins on contint1003 # [[phab:T436812|T436812]]
* 06:11 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet
* 06:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts2002.codfw.wmnet
* 06:00 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet
* 05:56 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts2002.codfw.wmnet
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 48s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:16 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1017.eqiad.wmnet with OS bookworm
== 2026-09-02 ==
* 23:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] (duration: 10m 21s)
* 23:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:50 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 23:49 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]]
* 23:38 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:38 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host aqs1017.eqiad.wmnet with OS bookworm
* 22:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] (duration: 11m 03s)
* 22:32 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 22:30 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:26 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]]
* 22:05 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] (duration: 14m 14s)
* 21:59 krinkle@deploy1003: krinkle: Continuing with deployment
* 21:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:50 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]]
* 21:44 inflatador: bking@apt1002 sudo -E private_reprepro --ignore=wrongdistribution -C matomo_plugins include bookworm-wikimedia-private matomo-plugin-customreports_5.5.0-1_amd64.changes [[phab:T431608|T431608]]
* 21:40 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] (duration: 09m 48s)
* 21:35 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:34 jforrester@deploy1003: jforrester: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:34 inflatador: bking@apt1002 sudo -E reprepro -C main include bookworm-wikimedia matomo-plugin-marketingcampaignsreporting_5.2.2-3_amd64.changes [[phab:T431608|T431608]]
* 21:30 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]]
* 21:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 21:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:03 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 20:50 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 20:27 dancy@deploy1003: Finished scap sync-world: testing (duration: 09m 21s)
* 20:18 dancy@deploy1003: Started scap sync-world: testing
* 20:18 dancy@deploy1003: Installation of scap version "4.288.0" completed for 3 hosts
* 20:16 dancy@deploy1003: Installing scap version "4.288.0" for 3 host(s)
* 19:57 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] (duration: 64m 27s)
* 19:55 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 18:57 jforrester@deploy1003: jforrester: Continuing with deployment
* 18:57 jforrester@deploy1003: jforrester: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:55 swfrench-wmf: deleted pods coredns-85b4f68d95-pk5sn coredns-85b4f68d95-22ddb coredns-85b4f68d95-49k5p in eqiad due to intermittent upstream resolution health check failures correlated with high DNS resolution latency
* 18:53 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]]
* 18:31 sukhe@dns1004: END - running authdns-update
* 18:28 sukhe@dns1004: START - running authdns-update
* 18:26 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:26 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:17 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:16 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 18:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:14 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:14 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqiad
* 18:14 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:49 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqiad
* 17:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:45 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqsin
* 17:38 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:38 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:20 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqsin
* 17:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 17:00 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 16:59 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:48 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 16:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:41 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 16:41 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 16:41 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:38 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:34 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 16:29 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_drmrs
* 16:24 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 16:14 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_drmrs
* 16:11 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-unlock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 16:10 root@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]] (duration: 48m 09s)
* 16:10 root@deploy1003: Forcefully removing global lock: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 16:10 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-unlock-scap for datacenter switchover from codfw to eqiad
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:59 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:58 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters for datacenter switchover from codfw to eqiad
* 15:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 15:58 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-restore-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-restore-ttl for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-start-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 root@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-start-maintenance for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-jobrunner: sync
* 15:56 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-jobrunner: sync
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.07-set-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period ends at: 2026-09-02 15:56:13.434320
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.07-set-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.03-set-db-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.03-set-db-readonly for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.02-set-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period starts at: 2026-09-02 15:53:47.690918
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.02-set-readonly for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.01-stop-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.01-stop-maintenance for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-reduce-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:47 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-reduce-ttl for datacenter switchover from codfw to eqiad
* 15:46 slyngshede@cumin1003: END (ERROR) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=97) for datacenter switchover from codfw to eqiad
* 15:45 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 15:44 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 15:42 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service
* 15:39 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:38 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:31 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:28 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:27 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 15:27 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from codfw to eqiad
* 15:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-lock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:22 root@deploy1003: Locking from deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-lock-scap for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks for datacenter switchover from codfw to eqiad
* 15:17 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service
* 15:15 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_magru
* 15:11 moritzm: import jenkins 2.568.3 to thirdparty/jenkins for trixie-wikimedia
* 14:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 14:44 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_magru
* 14:32 moritzm: installing pdns-recursor security updates
* 14:27 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:27 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:20 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:15 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 14:12 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 14:09 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:09 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:09 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:08 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:06 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:44 moritzm: bounce tcpircbot-logmsgbot/tcpircbot-logmsgbot_cloud on alert1002 to allow cumin1004 [[phab:T427897|T427897]]
* 13:36 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] (duration: 09m 52s)
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Continuing with deployment
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:26 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]]
* 13:25 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 13:24 moritzm: installing wireshark security updates
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 13:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:04 moritzm: import librsvg 2.60.0+dfsg-1+wmf13u1 to component/thumbor for trixie-wikimedia [[phab:T436505|T436505]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 12:44 atsuko@dns1004: END - running authdns-update
* 12:41 atsuko@dns1004: START - running authdns-update
* 12:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:35 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] (duration: 12m 50s)
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:24 dreamyjazz@deploy1003: dreamyjazz, btullis: Continuing with deployment
* 12:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:22 dreamyjazz@deploy1003: dreamyjazz, btullis: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:20 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:17 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]]
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 11:31 marostegui@cumin1003: Removing db1172 from zarcillo [[phab:T436763|T436763]]
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1172.eqiad.wmnet
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:30 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:26 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:26 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:25 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 11:25 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:24 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:20 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1172.eqiad.wmnet
* 11:20 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 11:12 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:11 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:10 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:09 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:05 slyngshede@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.*
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 11:05 slyngshede@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 11:03 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply
* 11:03 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
* 10:52 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:49 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:48 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Pool back db1242
* 10:45 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow4003.ulsfo.wmnet with OS trixie
* 10:31 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1172 from dbctl [[phab:T436763|T436763]]', diff saved to https://phabricator.wikimedia.org/P96318 and previous config saved to /var/cache/conftool/dbconfig/20260902-103152-marostegui.json
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 blake@deploy1003: Finished scap sync-world: non-build deployment for [[phab:T417800|T417800]] (duration: 05m 37s)
* 10:06 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:05 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:04 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:03 blake@deploy1003: Started scap sync-world: non-build deployment for [[phab:T417800|T417800]]
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 10:00 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:00 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:59 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:58 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:57 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:56 jmm@dns1004: END - running authdns-update
* 09:56 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:55 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:54 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:53 jmm@dns1004: START - running authdns-update
* 09:51 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:47 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1228 to dbctl [[phab:T435892|T435892]]', diff saved to https://phabricator.wikimedia.org/P96313 and previous config saved to /var/cache/conftool/dbconfig/20260902-094713-marostegui.json
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:42 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow4003.ulsfo.wmnet with OS trixie
* 09:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow7002.magru.wmnet with OS trixie
* 09:31 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:30 moritzm: installing openjdk-21 security updates
* 09:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:22 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:17 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:17 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:16 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:10 moritzm: installing openjdk-8 security updates
* 09:09 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:08 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:46 tappof: bump space for prometheus k8s-dse in eqiad
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:39 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow7002.magru.wmnet with OS trixie
* 08:36 moritzm: installing libgraphite2 security updates
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:23 Msz2001: UTC morning backport window done
* 08:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] (duration: 14m 36s)
* 08:19 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:19 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet
* 08:14 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:14 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet
* 08:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet
* 08:08 fabfur@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cp5022.eqsin.wmnet with reason: investigating
* 08:07 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]]
* 08:07 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp5022.*
* 08:07 fabfur: depooling and silencing cp5022 ([[phab:T414411|T414411]])
* 08:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet
* 08:03 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:03 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* {{safesubst:SAL entry|1=08:03 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)}}
* 07:49 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 07:49 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]] synced to the
* 07:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1004.eqiad.wmnet
* 07:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1004.eqiad.wmnet
* 07:30 jmm@dns1004: END - running authdns-update
* {{safesubst:SAL entry|1=07:27 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]}}
* 07:27 jmm@dns1004: START - running authdns-update
* 07:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] (duration: 16m 04s)
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1242: Cloning db1228
* 07:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1242: Cloning db1228
* 07:18 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1228,1242].eqiad.wmnet with reason: db1242 needs to clone db1228
* 07:17 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Continuing with deployment
* 07:12 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1228.eqiad.wmnet with OS trixie
* 07:10 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be veri
* 07:06 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]]
* 06:43 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: build: Updating npm dependencies (duration: 00m 13s)
* 06:43 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: build: Updating npm dependencies
* 06:39 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:32 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:23 slyngshede@dns1004: END - running authdns-update
* 06:21 marostegui: Drop cu* tables from s3 bswiktionary [[phab:T435965|T435965]]
* 06:20 slyngshede@dns1004: START - running authdns-update
* 06:18 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1228.eqiad.wmnet with OS trixie
* 06:13 XioNoX: re-enable magru cr1/asw1-b3 link - [[phab:T436675|T436675]]
* 05:06 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] (duration: 04m 42s)
* 05:04 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:03 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:01 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 05:01 tstarling@deploy1003: Scap cancelled without rolling back.
* 04:53 tstarling@deploy1003: tstarling: Continuing with deployment
* 04:29 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 04:25 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 43s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-01 ==
* 21:59 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] (duration: 18m 05s)
* 21:52 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:47 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]]
* 21:38 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] (duration: 23m 55s)
* 21:28 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:20 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:14 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]]
* 20:47 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1016.eqiad.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:24 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:11 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 20:11 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:57 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:57 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:56 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:54 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:52 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:47 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:45 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:42 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:41 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:40 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:39 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:32 jhancock@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudcephosd1055
* 19:32 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:30 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.* [reason: update IP addrs]
* 19:30 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 19:30 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:23 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:22 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:16 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:13 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1056
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1056
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1055
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:12 jclark@cumin1003: END (ERROR) - Cookbook sre.network.configure-switch-interfaces (exit_code=97) for host cloudcephosd1054
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1054
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:11 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:06 jclark@cumin1003: START - Cookbook sre.dns.netbox
* 19:06 sukhe@dns1004: END - running authdns-update
* 19:03 sukhe@dns1004: START - running authdns-update
* 18:14 dancy@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 17:04 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 17:00 dancy@deploy1003: Finished scap sync-world: testing (duration: 08m 07s)
* 16:52 dancy@deploy1003: Started scap sync-world: testing
* 16:48 dancy@deploy1003: sync-world aborted: testing (duration: 00m 05s)
* 16:48 dancy@deploy1003: Started scap sync-world: testing
* 16:47 dancy@deploy1003: Installation of scap version "4.287.0" completed for 156 hosts
* 16:42 dancy@deploy1003: Installing scap version "4.287.0" for 156 host(s)
* 16:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 16:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 15:51 moritzm: installing mesa security updates
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts phab1004.eqiad.wmnet
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:27 aokoth@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:20 aokoth@cumin1003: START - Cookbook sre.dns.netbox
* 15:14 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595] (duration: 05m 32s)
* 15:14 aokoth@cumin1003: START - Cookbook sre.hosts.decommission for hosts phab1004.eqiad.wmnet
* 15:11 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on phab1004.eqiad.wmnet with reason: Decom
* 15:09 joal@deploy1003: Started deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595]
* 15:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config: apply
* 15:08 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config: apply
* 14:55 hashar: Restarted Jenkins on releases1003
* 14:51 hashar: Restarted CI Jenkins on contint1003
* 14:48 hashar: Restarting Gerrit primary on gerrit2003
* 14:45 hashar: Restarted Gerrit on gerrit1003 and gerrit2002
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 14:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 14:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 14:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:24 moritzm: installing curl security updates
* 14:24 jmm@dns1004: END - running authdns-update
* 14:23 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser (duration: 00m 14s)
* 14:23 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:21 jmm@dns1004: START - running authdns-update
* 14:21 jmm@dns1004: END - running authdns-update
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 14:19 jmm@dns1004: START - running authdns-update
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet
* 14:14 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:14 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:13 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595] (duration: 07m 26s)
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Test
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Test
* 14:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet
* 14:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Test
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Test
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 14:05 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595]
* 14:05 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] (duration: 37m 37s)
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 hashar: Removed openjdk-17 packages from contint1002/contint2002 following relocation of CI Jenkins to contint1003/contint2003 # [[phab:T418521|T418521]]
* 14:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:02 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:02 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:00 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595] (duration: 00m 45s)
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 13:59 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595]
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:58 ladsgroup@dns1004: END - running authdns-update
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:57 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 ladsgroup@dns1004: START - running authdns-update
* 13:56 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:56 ladsgroup@dns1004: END - running authdns-update
* 13:55 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:53 ladsgroup@dns1004: START - running authdns-update
* 13:50 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:49 kharlan@deploy1003: kharlan: Continuing with deployment
* 13:48 kharlan@deploy1003: kharlan: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2004.wikimedia.org
* 13:42 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:41 jmm@dns1004: END - running authdns-update
* 13:40 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2004.wikimedia.org
* 13:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2003.wikimedia.org
* 13:38 jmm@dns1004: START - running authdns-update
* 13:34 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2003.wikimedia.org
* 13:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1004.wikimedia.org
* 13:26 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]]
* 13:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:25 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:24 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1004.wikimedia.org
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:23 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] (duration: 20m 24s)
* 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1003.wikimedia.org
* 13:20 fnegri@deploy1003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply
* 13:19 fnegri@deploy1003: helmfile [eqiad] START helmfile.d/services/toolhub: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:16 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1003.wikimedia.org
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:15 moritzm: bump urldownloader[12]00[34] to 8G RAM [[phab:T429175|T429175]]
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:15 fnegri@deploy1003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 fnegri@deploy1003: helmfile [codfw] START helmfile.d/services/toolhub: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:11 fnegri@deploy1003: helmfile [staging] DONE helmfile.d/services/toolhub: apply
* 13:11 aude@deploy1003: aude: Continuing with deployment
* 13:10 fnegri@deploy1003: helmfile [staging] START helmfile.d/services/toolhub: apply
* 13:07 aude@deploy1003: aude: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2004.wikimedia.org with OS bookworm
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow1004.eqiad.wmnet
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1004.eqiad.wmnet with OS trixie
* 12:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:47 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:41 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] (duration: 16m 25s)
* 12:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 12:34 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 12:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 12:33 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 12:31 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:29 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:25 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]]
* 12:22 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1004.eqiad.wmnet with OS trixie
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:20 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 12:20 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 12:16 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow1004.eqiad.wmnet
* 12:15 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2004.wikimedia.org with OS bookworm
* 12:14 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 12:07 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2003.wikimedia.org with OS bookworm
* 11:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:43 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:32 jmm@cumin2003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling restart_daemons on A:kafka-test-eqiad
* 11:26 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2003.wikimedia.org with OS bookworm
* 11:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1004.wikimedia.org with OS bookworm
* 11:12 moritzm: installing openjdk-21 security updates
* 11:12 jmm@cumin2003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling restart_daemons on A:kafka-test-eqiad
* 10:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:53 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:49 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2902: Pool back db2902
* 10:45 moritzm: installing Python 3.11 security updates
* 10:37 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1004.wikimedia.org with OS bookworm
* 10:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5022.eqsin.wmnet with OS trixie
* 10:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp5022.eqsin.wmnet on all recursors
* 10:36 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache cp5022.eqsin.wmnet on all recursors
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:30 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:04 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:04 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2902: Pool back db2902
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:03 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2902: test
* 10:03 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 10:01 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 09:58 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:50 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/echoserver: apply
* 09:49 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/echoserver: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Change db1176 and db2230's weight, test-s4 masters, to 0 to mimic the rest of production [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96292 and previous config saved to /var/cache/conftool/dbconfig/20260901-092444-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96291 and previous config saved to /var/cache/conftool/dbconfig/20260901-090233-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96290 and previous config saved to /var/cache/conftool/dbconfig/20260901-090158-marostegui.json
* 09:01 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96289 and previous config saved to /var/cache/conftool/dbconfig/20260901-090121-marostegui.json
* 09:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:56 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:50 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1003.wikimedia.org with OS bookworm
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1074.eqiad.wmnet
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:46 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:43 marostegui@cumin1003: dbctl commit (dc=all): 'Test repool db2902', diff saved to https://phabricator.wikimedia.org/P96288 and previous config saved to /var/cache/conftool/dbconfig/20260901-084317-marostegui.json
* 08:42 marostegui@cumin1003: dbctl commit (dc=all): 'Test depool db2902', diff saved to https://phabricator.wikimedia.org/P96287 and previous config saved to /var/cache/conftool/dbconfig/20260901-084249-marostegui.json
* 08:39 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:37 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2902: test
* 08:36 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96285 and previous config saved to /var/cache/conftool/dbconfig/20260901-083557-marostegui.json
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96284 and previous config saved to /var/cache/conftool/dbconfig/20260901-083527-marostegui.json
* 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96283 and previous config saved to /var/cache/conftool/dbconfig/20260901-083432-marostegui.json
* 08:32 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1074.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1073.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:27 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:24 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 08:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022']
* 08:16 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1073.eqiad.wmnet
* 08:15 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1003.wikimedia.org with OS bookworm
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1072.eqiad.wmnet
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:14 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:10 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:05 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1072.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1067.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:00 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:56 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:54 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022']
* 07:53 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 07:52 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie
* 07:52 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 07:50 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1067.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1066.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:47 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 07:41 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:34 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1066.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1065.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:32 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:27 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:23 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1065.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1075.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:15 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 06:49 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:45 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1075.eqiad.wmnet
* 06:29 moritzm: installing Java 17 security updates
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.15 (duration: 02m 25s)
* 03:50 denisse@deploy1003: Finished deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2 (duration: 00m 19s)
* 03:50 denisse@deploy1003: Started deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2
* 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]] (duration: 37m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:30 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 00:29 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
== Other archives ==
See [[Server Admin Log/Archives]].
<noinclude>
[[Category:SAL]]
[[Category:Operations]]
</noinclude>
2kccjq6a7hv4sq70i8h00n3t511ykp9
2458701
2458699
2026-09-19T16:55:27Z
Stashbot
7414
ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
2458701
wikitext
text/x-wiki
== 2026-09-19 ==
* 16:55 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:11 urbanecm: Attach SHB@commonswiki to the SUL account manually ([[phab:T438591|T438591]], see [[phab:T438591|T438591]]#12341750 for what I did exactly)
* 04:08 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 04:08 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 04:08 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 04:07 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 36s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-18 ==
* 22:41 rzl@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=sessionstore,name=eqiad
* 17:08 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343065{{!}}mw.DesktopArticleTarget: if source education is enabled suppress welcome (T434249)]] (duration: 09m 26s)
* 17:05 Dreamy_Jazz: Created `securepoll_log` on `nlwiki` main DB cluster for [[phab:T434045|T434045]]
* 17:04 kemayo@deploy1003: kemayo: Continuing with deployment
* 17:03 kemayo@deploy1003: kemayo: Backport for [[gerrit:1343065{{!}}mw.DesktopArticleTarget: if source education is enabled suppress welcome (T434249)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:59 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1343065{{!}}mw.DesktopArticleTarget: if source education is enabled suppress welcome (T434249)]]
* 16:49 oblivian@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343127{{!}}ResourceLoader: hotfix for current logspam over the weekend (T438387)]] (duration: 11m 36s)
* 16:42 oblivian@deploy1003: oblivian: Continuing with deployment
* 16:42 oblivian@deploy1003: oblivian: Backport for [[gerrit:1343127{{!}}ResourceLoader: hotfix for current logspam over the weekend (T438387)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:37 oblivian@deploy1003: Started scap sync-world: Backport for [[gerrit:1343127{{!}}ResourceLoader: hotfix for current logspam over the weekend (T438387)]]
* 16:09 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ml-serve1016.eqiad.wmnet with OS trixie
* 14:49 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie
* 13:37 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host registry1004.eqiad.wmnet with OS trixie
* 13:23 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on registry1004.eqiad.wmnet with reason: host reimage
* 13:18 elukey@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on registry1004.eqiad.wmnet with reason: host reimage
* 13:04 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host registry1004.eqiad.wmnet with OS trixie
* 12:14 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:13 jclark@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:30 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 09:30 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 09:22 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on an-redacteddb1001.eqiad.wmnet with reason: cloning
* 09:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 09:19 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 09:18 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 21 hosts with reason: cloning db1270
* 09:18 marostegui: clone db1270:x4 from db1155:x4 lag will appear on x4
* 09:09 brouberol@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 09:08 brouberol@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 08:23 brouberol@dns1004: END - running authdns-update
* 08:21 brouberol@dns1004: START - running authdns-update
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 57s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-17 ==
* 21:04 tsev@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T438395|T438395]]
* 20:55 tsev@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T438395|T438395]]
* 20:47 jhuneidi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342774{{!}}Worklist Promotion test kitchen - Enable flag in production (T434513)]], [[gerrit:1342781{{!}}Exclude returntoapp query from app interception on iOS (T438395)]], [[gerrit:1342798{{!}}Revert "Update wikimania wordmark for 2026"]] (duration: 35m 59s)
* 20:35 jhuneidi@deploy1003: robertsky, jhuneidi, cmelo, tsev: Continuing with deployment
* 20:31 jhuneidi@deploy1003: robertsky, jhuneidi, cmelo, tsev: Backport for [[gerrit:1342774{{!}}Worklist Promotion test kitchen - Enable flag in production (T434513)]], [[gerrit:1342781{{!}}Exclude returntoapp query from app interception on iOS (T438395)]], [[gerrit:1342798{{!}}Revert "Update wikimania wordmark for 2026"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:11 jhuneidi@deploy1003: Started scap sync-world: Backport for [[gerrit:1342774{{!}}Worklist Promotion test kitchen - Enable flag in production (T434513)]], [[gerrit:1342781{{!}}Exclude returntoapp query from app interception on iOS (T438395)]], [[gerrit:1342798{{!}}Revert "Update wikimania wordmark for 2026"]]
* 19:15 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1100.eqiad.wmnet with OS trixie
* 19:15 vriley@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1004"
* 19:10 vriley@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1004"
* 18:52 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1100.eqiad.wmnet with reason: host reimage
* 18:48 vriley@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1100.eqiad.wmnet with reason: host reimage
* 18:32 vriley@cumin1004: START - Cookbook sre.hosts.reimage for host ms-be1100.eqiad.wmnet with OS trixie
* 18:20 urbanecm: Deploy a security fix for [[phab:T438389|T438389]]
* 17:47 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be1100.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:38 vriley@cumin1004: START - Cookbook sre.hosts.provision for host ms-be1100.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:38 vriley@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1100.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:37 vriley@cumin1004: START - Cookbook sre.hosts.provision for host ms-be1100.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:23 vriley@cumin1004: START - Cookbook sre.hosts.reimage for host ms-be1100.eqiad.wmnet with OS trixie
* 16:52 aokoth@deploy1003: Finished deploy [phabricator/deployment@c386249]: Deploy Phab (duration: 00m 12s)
* 16:52 aokoth@deploy1003: Started deploy [phabricator/deployment@c386249]: Deploy Phab
* 16:42 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1099.eqiad.wmnet with OS trixie
* 16:42 vriley@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1004"
* 16:42 vriley@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1004"
* 16:35 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be1100.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:31 aokoth@deploy1003: Finished deploy [phabricator/deployment@c386249]: Deploy Phab (duration: 00m 19s)
* 16:31 aokoth@deploy1003: Started deploy [phabricator/deployment@c386249]: Deploy Phab
* 16:21 vriley@cumin1004: START - Cookbook sre.hosts.provision for host ms-be1100.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:20 vriley@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be1100
* 16:20 vriley@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host ms-be1100
* 16:19 vriley@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:19 vriley@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [ms-be1100] - vriley@cumin1004"
* 16:19 vriley@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [ms-be1100] - vriley@cumin1004"
* 16:15 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 16:15 vriley@cumin1004: START - Cookbook sre.dns.netbox
* 16:15 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 16:14 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 16:14 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 16:12 moritzm: installing libapache-mod-auth-oidc security updates
* 16:12 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1099.eqiad.wmnet with reason: host reimage
* 16:11 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 16:11 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 16:08 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/editcheck-headless: apply
* 16:07 rzl@deploy1003: helmfile [staging] START helmfile.d/services/editcheck-headless: apply
* 16:06 vriley@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1099.eqiad.wmnet with reason: host reimage
* 16:01 moritzm: installing aom security updates
* 16:01 btullis@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ceph-admin2001.codfw.wmnet
* 16:01 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ceph-admin2001.codfw.wmnet with OS bookworm
* 15:55 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 15:55 rzl@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 15:54 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 15:54 rzl@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 15:54 rzl@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 15:53 rzl@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 15:53 rzl@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 15:52 rzl@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 15:51 vriley@cumin1004: START - Cookbook sre.hosts.reimage for host ms-be1099.eqiad.wmnet with OS trixie
* 15:48 moritzm: installing libde265 security updates
* 15:44 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ceph-admin2001.codfw.wmnet with reason: host reimage
* 15:40 btullis@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ceph-admin1001.eqiad.wmnet
* 15:40 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ceph-admin1001.eqiad.wmnet with OS bookworm
* 15:39 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir4003.*
* 15:37 btullis@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ceph-admin2001.codfw.wmnet with reason: host reimage
* 15:26 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir4003.ulsfo.wmnet with OS trixie
* 15:25 aokoth@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host phab2003.codfw.wmnet with OS trixie
* 15:23 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ceph-admin1001.eqiad.wmnet with reason: host reimage
* 15:19 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host registry1005.eqiad.wmnet with OS trixie
* 15:17 btullis@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ceph-admin1001.eqiad.wmnet with reason: host reimage
* 15:16 btullis@cumin1004: START - Cookbook sre.hosts.reimage for host ceph-admin2001.codfw.wmnet with OS bookworm
* 15:16 btullis@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ceph-admin2001.codfw.wmnet - btullis@cumin1004"
* 15:16 btullis@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ceph-admin2001.codfw.wmnet - btullis@cumin1004"
* 15:15 btullis@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ceph-admin2001.codfw.wmnet on all recursors
* 15:15 btullis@cumin1004: START - Cookbook sre.dns.wipe-cache ceph-admin2001.codfw.wmnet on all recursors
* 15:15 btullis@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:15 btullis@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ceph-admin2001.codfw.wmnet - btullis@cumin1004"
* 15:15 btullis@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ceph-admin2001.codfw.wmnet - btullis@cumin1004"
* 15:08 aokoth@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on phab2003.codfw.wmnet with reason: host reimage
* 15:06 Msz2001: Deployed private code changes to Suggestedinvestigations
* 15:05 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir4003.ulsfo.wmnet with reason: host reimage
* 15:05 aokoth@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on phab2003.codfw.wmnet with reason: host reimage
* 15:04 btullis@cumin1004: START - Cookbook sre.hosts.reimage for host ceph-admin1001.eqiad.wmnet with OS bookworm
* 15:04 btullis@cumin1004: START - Cookbook sre.dns.netbox
* 15:04 btullis@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ceph-admin1001.eqiad.wmnet - btullis@cumin1004"
* 15:04 btullis@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ceph-admin1001.eqiad.wmnet - btullis@cumin1004"
* 15:04 btullis@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ceph-admin1001.eqiad.wmnet on all recursors
* 15:04 btullis@cumin1004: START - Cookbook sre.dns.wipe-cache ceph-admin1001.eqiad.wmnet on all recursors
* 15:04 btullis@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:04 btullis@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ceph-admin1001.eqiad.wmnet - btullis@cumin1004"
* 15:04 btullis@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ceph-admin1001.eqiad.wmnet - btullis@cumin1004"
* 15:02 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on registry1005.eqiad.wmnet with reason: host reimage
* 15:01 btullis@cumin1004: START - Cookbook sre.ganeti.makevm for new host ceph-admin2001.codfw.wmnet
* 15:00 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342265{{!}}JsonSchemaBuilder: Cache the root schema in the process (T437588)]], [[gerrit:1342264{{!}}JsonSchemaBuilder: Cache the root schema in the process (T437588)]], [[gerrit:1342684{{!}}SI: Preserve the username filter when switching queues (T438308)]] (duration: 12m 34s)
* 15:00 btullis@cumin1004: START - Cookbook sre.dns.netbox
* 15:00 btullis@cumin1004: START - Cookbook sre.ganeti.makevm for new host ceph-admin1001.eqiad.wmnet
* 14:58 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir4003.ulsfo.wmnet with reason: host reimage
* 14:58 elukey@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on registry1005.eqiad.wmnet with reason: host reimage
* 14:55 urbanecm@deploy1003: mszwarc, urbanecm: Continuing with deployment
* 14:51 urbanecm@deploy1003: mszwarc, urbanecm: Backport for [[gerrit:1342265{{!}}JsonSchemaBuilder: Cache the root schema in the process (T437588)]], [[gerrit:1342264{{!}}JsonSchemaBuilder: Cache the root schema in the process (T437588)]], [[gerrit:1342684{{!}}SI: Preserve the username filter when switching queues (T438308)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:51 aokoth@cumin1004: START - Cookbook sre.hosts.reimage for host phab2003.codfw.wmnet with OS trixie
* 14:50 aokoth@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on phab2003.codfw.wmnet with reason: Reimage
* 14:47 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1342265{{!}}JsonSchemaBuilder: Cache the root schema in the process (T437588)]], [[gerrit:1342264{{!}}JsonSchemaBuilder: Cache the root schema in the process (T437588)]], [[gerrit:1342684{{!}}SI: Preserve the username filter when switching queues (T438308)]]
* 14:42 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1331830{{!}}Keep Balinese Palm Leaf variants enabled on wikisource (T436398)]], [[gerrit:1340216{{!}}Turn on variant conversion for PageAssessments (T328012)]], [[gerrit:1341949{{!}}Parsoid Read Views: Enable on 61 wikiquote wikis (T437917)]] (duration: 15m 31s)
* 14:39 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host registry1005.eqiad.wmnet with OS trixie
* 14:35 cscott@deploy1003: ssastry, cscott: Continuing with deployment
* 14:33 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir4003.ulsfo.wmnet with OS trixie
* 14:32 cscott@deploy1003: ssastry, cscott: Backport for [[gerrit:1331830{{!}}Keep Balinese Palm Leaf variants enabled on wikisource (T436398)]], [[gerrit:1340216{{!}}Turn on variant conversion for PageAssessments (T328012)]], [[gerrit:1341949{{!}}Parsoid Read Views: Enable on 61 wikiquote wikis (T437917)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:26 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1331830{{!}}Keep Balinese Palm Leaf variants enabled on wikisource (T436398)]], [[gerrit:1340216{{!}}Turn on variant conversion for PageAssessments (T328012)]], [[gerrit:1341949{{!}}Parsoid Read Views: Enable on 61 wikiquote wikis (T437917)]]
* 14:20 caro@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342678{{!}}enwiki desktop VE: add education popup for switching to source editor (T434249)]], [[gerrit:1342362{{!}}Make VE the default editor on enwiki desktop (T436574)]] (duration: 33m 52s)
* 14:07 caro@deploy1003: caro: Continuing with deployment
* 14:06 caro@deploy1003: caro: Backport for [[gerrit:1342678{{!}}enwiki desktop VE: add education popup for switching to source editor (T434249)]], [[gerrit:1342362{{!}}Make VE the default editor on enwiki desktop (T436574)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:46 caro@deploy1003: Started scap sync-world: Backport for [[gerrit:1342678{{!}}enwiki desktop VE: add education popup for switching to source editor (T434249)]], [[gerrit:1342362{{!}}Make VE the default editor on enwiki desktop (T436574)]]
* 13:37 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 13:34 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:33 elukey@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:31 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:30 elukey@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:29 Emperor: apus - radosgw-admin quota set --quota-scope=user --uid=docker-registry --max-size=5T [[phab:T438339|T438339]]
* 13:24 jclark@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1016
* 13:24 jclark@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1016
* 13:11 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:11 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:09 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:09 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:55 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:53 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 12:42 jclark@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:40 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:40 jclark@cumin1003: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:35 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 12:19 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342626{{!}}MathMathML: Simplify Mathoid fallback/a11y class logic (T436026)]], [[gerrit:1342627{{!}}ext.math.mathjax: Implement mwe-math-mathml-a11y for client-side MathJax (T436026)]] (duration: 13m 04s)
* 12:14 krinkle@deploy1003: krinkle: Continuing with deployment
* 12:10 krinkle@deploy1003: krinkle: Backport for [[gerrit:1342626{{!}}MathMathML: Simplify Mathoid fallback/a11y class logic (T436026)]], [[gerrit:1342627{{!}}ext.math.mathjax: Implement mwe-math-mathml-a11y for client-side MathJax (T436026)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:05 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1342626{{!}}MathMathML: Simplify Mathoid fallback/a11y class logic (T436026)]], [[gerrit:1342627{{!}}ext.math.mathjax: Implement mwe-math-mathml-a11y for client-side MathJax (T436026)]]
* 10:37 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 10:28 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 10:10 blake@deploy1003: Finished scap sync-world: cleanup for [[phab:T417800|T417800]] (duration: 03m 57s)
* 10:07 blake@deploy1003: Started scap sync-world: cleanup for [[phab:T417800|T417800]]
* 09:52 marostegui@cumin1004: dbctl commit (dc=all): 'Fix weights [[phab:T436496|T436496]]', diff saved to https://phabricator.wikimedia.org/P96466 and previous config saved to /var/cache/conftool/dbconfig/20260917-095235-marostegui.json
* 09:51 marostegui@cumin1004: dbctl commit (dc=all): 'Fix weights [[phab:T436496|T436496]]', diff saved to https://phabricator.wikimedia.org/P96465 and previous config saved to /var/cache/conftool/dbconfig/20260917-095131-marostegui.json
* 09:41 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:41 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 09:41 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:40 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 09:40 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:40 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 09:35 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 09:35 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 08:37 moritzm: pruned obsolete Bullseye image prometheus-nutcracker-exporter from the docker registry [[phab:T416452|T416452]]
* 08:34 XioNoX: Manually install gnmic 0.49.0 on netflow2005 - [[phab:T438291|T438291]]
* 08:28 brouberol@dns1004: END - running authdns-update
* 08:26 brouberol@dns1004: START - running authdns-update
* 08:13 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.20 refs [[phab:T430839|T430839]]
* 08:10 moritzm: imported nodejs_26.8.2-1nodesource1 to thirdparty/node26 for trixie-wikimedia [[phab:T437510|T437510]]
* 08:07 Amir1: dropped links tables from db2206 ([[phab:T437278|T437278]])
* 08:03 Amir1: dropped links tables from db2219 ([[phab:T437278|T437278]])
* 08:01 Amir1: dropped links tables from db2236 ([[phab:T437278|T437278]])
* 07:59 Amir1: dropped non-links tables from db1262 ([[phab:T437278|T437278]])
* 07:57 Amir1: dropped non-links tables from db2245 ([[phab:T437278|T437278]])
* 07:52 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@8eaca67] (releasing): [[phab:T438205|T438205]] to prod host (duration: 00m 44s)
* 07:52 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@8eaca67] (releasing): [[phab:T438205|T438205]] to prod host
* 07:49 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@8eaca67] (releasing): [[phab:T438205|T438205]] to backup host (duration: 00m 47s)
* 07:48 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@8eaca67] (releasing): [[phab:T438205|T438205]] to backup host
* 07:25 XioNoX: Manually install gnmic 0.49.0 on netflow1004 - [[phab:T438291|T438291]]
* 07:24 mlitn@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342398{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342396{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342395{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342545{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342546{{!}}Localisation updates from https://translatewiki.net.]],
* 07:19 mlitn@deploy1003: mlitn, jdlrobson: Continuing with deployment
* {{safesubst:SAL entry|1=07:18 mlitn@deploy1003: mlitn, jdlrobson: Backport for [[gerrit:1342398{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342396{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342395{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342545{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342546{{!}}Localisation updates from https://translatewiki.net.]], [[gerri}}
* 07:11 mlitn@deploy1003: Started scap sync-world: Backport for [[gerrit:1342398{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342396{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342395{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342545{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342546{{!}}Localisation updates from https://translatewiki.net.]],
* 07:10 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Jmoore111 out of all services on: 2444 hosts
* 06:07 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 05:53 marostegui@cumin1004: END (FAIL) - Cookbook sre.mysql.decommission (exit_code=99)
* 05:53 marostegui@cumin1004: Removing db1180 from zarcillo [[phab:T437222|T437222]]
* 05:53 marostegui@cumin1004: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1180.eqiad.wmnet
* 05:53 marostegui@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 05:53 marostegui@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1180.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1004"
* 05:53 marostegui@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1180.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1004"
* 05:49 marostegui@cumin1004: START - Cookbook sre.dns.netbox
* 05:44 marostegui@cumin1004: START - Cookbook sre.hosts.decommission for hosts db1180.eqiad.wmnet
* 05:43 marostegui@cumin1004: START - Cookbook sre.mysql.decommission
* 04:26 aokoth@cumin1004: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet
* 04:24 aokoth@cumin1004: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet
== 2026-09-16 ==
* 23:10 rzl: rzl@deploy1003 Finished scap sync-world: Backport for [[gerrit:1342091{{!}}Repool poolcounter[1007,2006] (T435163)]] (duration: 11m 09s)
* 22:50 rzl@deploy1003: Started scap sync-world: Backport for [[gerrit:1342091{{!}}Repool poolcounter[1007,2006] (T435163)]]
* 22:47 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342383{{!}}DonorIdentification: Confirm before unlinking donor status in preferences (T436698)]], [[gerrit:1342385{{!}}Make learn more link to new window (T438252)]] (duration: 35m 21s)
* 22:35 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 22:33 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1342383{{!}}DonorIdentification: Confirm before unlinking donor status in preferences (T436698)]], [[gerrit:1342385{{!}}Make learn more link to new window (T438252)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:12 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1342383{{!}}DonorIdentification: Confirm before unlinking donor status in preferences (T436698)]], [[gerrit:1342385{{!}}Make learn more link to new window (T438252)]]
* 22:10 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2006.codfw.wmnet
* 22:06 rzl@cumin2003: START - Cookbook sre.hosts.reboot-single for host poolcounter2006.codfw.wmnet
* 22:06 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1007.eqiad.wmnet
* 22:02 rzl@cumin2003: START - Cookbook sre.hosts.reboot-single for host poolcounter1007.eqiad.wmnet
* 21:56 rzl@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342090{{!}}Repool poolcounter[1006,2005]; depool poolcounter[1007,2006] for reboot (T435163)]] (duration: 09m 39s)
* 21:52 rzl@deploy1003: rzl: Continuing with deployment
* 21:51 rzl@deploy1003: rzl: Backport for [[gerrit:1342090{{!}}Repool poolcounter[1006,2005]; depool poolcounter[1007,2006] for reboot (T435163)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:47 rzl@deploy1003: Started scap sync-world: Backport for [[gerrit:1342090{{!}}Repool poolcounter[1006,2005]; depool poolcounter[1007,2006] for reboot (T435163)]]
* 21:46 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:43 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2005.codfw.wmnet
* 21:42 vriley@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1099.eqiad.wmnet with OS trixie
* 21:41 vriley@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1098.eqiad.wmnet with OS trixie
* 21:41 vriley@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin2003"
* 21:40 vriley@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin2003"
* 21:40 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:40 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:39 rzl@cumin2003: START - Cookbook sre.hosts.reboot-single for host poolcounter2005.codfw.wmnet
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1006.eqiad.wmnet
* 21:38 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:38 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:37 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:35 rzl@cumin2003: START - Cookbook sre.hosts.reboot-single for host poolcounter1006.eqiad.wmnet
* 21:32 rzl@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342089{{!}}Depool poolcounter[1006,2005] for reboot (T435163)]] (duration: 13m 53s)
* 21:26 rzl@deploy1003: rzl: Continuing with deployment
* 21:25 rzl@deploy1003: rzl: Backport for [[gerrit:1342089{{!}}Depool poolcounter[1006,2005] for reboot (T435163)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:23 vriley@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1098.eqiad.wmnet with reason: host reimage
* 21:18 rzl@deploy1003: Started scap sync-world: Backport for [[gerrit:1342089{{!}}Depool poolcounter[1006,2005] for reboot (T435163)]]
* 21:17 vriley@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1098.eqiad.wmnet with reason: host reimage
* 21:10 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:09 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:09 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:09 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:08 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:02 vriley@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS trixie
* 20:49 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342339{{!}}Reapply "Tell VisualEditor about the app web edit tags", modified]] (duration: 35m 55s)
* 20:37 kemayo@deploy1003: cklimas, kemayo: Continuing with deployment
* 20:33 kemayo@deploy1003: cklimas, kemayo: Backport for [[gerrit:1342339{{!}}Reapply "Tell VisualEditor about the app web edit tags", modified]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:13 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1342339{{!}}Reapply "Tell VisualEditor about the app web edit tags", modified]]
* 19:20 dwisehaupt@dns1005: END - running authdns-update
* 19:18 dwisehaupt@dns1005: START - running authdns-update
* 19:06 vriley@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:57 dwisehaupt@dns1005: END - running authdns-update
* 18:55 vriley@cumin2003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:55 dwisehaupt@dns1005: START - running authdns-update
* 18:44 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 18:42 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 18:37 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 18:35 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 18:26 robh@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:22 robh@cumin2003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:21 dzahn@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 18:21 robh@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:21 robh@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:20 dzahn@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 18:20 dzahn@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 18:18 dzahn@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 18:18 mutante: k8s/miscweb: admin_ng deploy: creating namespace for attribution.wikimedia.org [[phab:T437635|T437635]]
* 18:17 dzahn@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 18:17 dzahn@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 18:17 dzahn@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 18:16 dzahn@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 18:13 cdanis@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "fix known-client creation - cdanis@cumin1003"
* 18:13 cdanis@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: fix known-client creation - cdanis@cumin1003
* 18:12 cdanis@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: fix known-client creation - cdanis@cumin1003
* 18:12 cdanis@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "fix known-client creation - cdanis@cumin1003"
* 18:04 vriley@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:51 vriley@cumin2003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:46 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 17:46 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 17:45 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:45 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:44 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:44 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:35 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1006.eqiad.wmnet with OS trixie
* 17:35 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1003"
* 17:30 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:29 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1003"
* 17:27 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:27 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:23 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:23 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:20 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 17:20 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 17:14 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:13 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1006.eqiad.wmnet with reason: host reimage
* 17:10 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:10 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:10 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:10 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:09 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:07 jclark@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1006.eqiad.wmnet with reason: host reimage
* 17:07 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 17:05 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:05 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:52 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie
* 16:46 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir7003.magru.wmnet
* 16:44 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir7003
* 16:19 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie
* 15:42 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply
* 15:42 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply
* 15:42 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 15:42 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir7003.magru.wmnet with OS trixie
* 15:41 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 15:36 jnuche@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342279{{!}}Use parser output value instead of status (T438154)]] (duration: 33m 21s)
* 15:35 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply
* 15:35 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply
* 15:24 jnuche@deploy1003: jnuche, jforrester: Continuing with deployment
* 15:23 jnuche@deploy1003: jnuche, jforrester: Backport for [[gerrit:1342279{{!}}Use parser output value instead of status (T438154)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:19 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie
* 15:18 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir7003.magru.wmnet with reason: host reimage
* 15:14 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir7003.magru.wmnet with reason: host reimage
* 15:03 jnuche@deploy1003: Started scap sync-world: Backport for [[gerrit:1342279{{!}}Use parser output value instead of status (T438154)]]
* 14:50 moritzm: installing apache2 security updates
* 14:49 slyngshede@cumin1003: conftool action : set/pooled=yes; selector: name=cp5026.eqsin.wmnet
* 14:47 slyngshede@cumin1003: conftool action : set/weight=1; selector: name=cp5026.eqsin.wmnet
* 14:45 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5026.eqsin.wmnet with OS trixie
* 14:44 moritzm: installing python-filelock security updates
* 14:42 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir7003.magru.wmnet with OS trixie
* 14:35 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:35 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:34 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:33 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp6002.drmrs.wmnet
* 14:32 sukhe@puppetserver1001: conftool action : set/weight=100; selector: name=cp6002.drmrs.wmnet,service=ats-be
* 14:32 sukhe@puppetserver1001: conftool action : set/weight=1; selector: name=cp6002.drmrs.wmnet,service=cdn
* 14:29 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6002.drmrs.wmnet with OS trixie
* 14:24 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: sync
* 14:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:24 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:24 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: sync
* 14:14 jforrester@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.19,1.47.0-wmf.20,next --multiversion-image-basename docker-registry.discovery.wmnet/restricte
* 14:14 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:14 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:13 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1342008{{!}}abstractwiki: Add three new articles per community advice to show off the feature (T434227)]]
* 14:10 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: sync
* 14:10 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: sync
* 14:10 jforrester@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.19,1.47.0-wmf.20,next --multiversion-image-basename docker-registry.discovery.wmnet/restricte
* 14:10 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: sync
* 14:10 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: sync
* 14:09 Amir1: dropped links tables on db2237 ([[phab:T437278|T437278]])
* 14:08 Amir1: dropped links tables on db1238 ([[phab:T437278|T437278]])
* 14:07 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1342008{{!}}abstractwiki: Add three new articles per community advice to show off the feature (T434227)]]
* 14:03 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5026.eqsin.wmnet with reason: host reimage
* 14:02 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:02 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:02 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 14:01 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:00 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 13:59 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6002.drmrs.wmnet with reason: host reimage
* 13:56 slyngshede@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5026.eqsin.wmnet with reason: host reimage
* 13:54 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cp6002.drmrs.wmnet with reason: host reimage
* 13:53 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 13:52 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 13:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dse-k8s-etcd[1001-1003].eqiad.wmnet
* 13:48 bking@cumin2003: START - Cookbook sre.hosts.remove-downtime for dse-k8s-etcd[1001-1003].eqiad.wmnet
* 13:46 bking@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM dse-k8s-etcd1001.eqiad.wmnet
* 13:46 bking@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM dse-k8s-etcd1001.eqiad.wmnet
* 13:45 bking@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM dse-k8s-etcd1002.eqiad.wmnet
* 13:41 bking@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM dse-k8s-etcd1002.eqiad.wmnet
* 13:41 bking@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM dse-k8s-etcd1003.eqiad.wmnet
* 13:38 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp6002.drmrs.wmnet with OS trixie
* 13:37 bking@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM dse-k8s-etcd1003.eqiad.wmnet
* 13:37 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.reboot-vm (exit_code=99) for VM dse-k8s-etcd1003.eqiad.wmnet
* 13:37 bking@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM dse-k8s-etcd1003.eqiad.wmnet
* 13:34 slyngshede@cumin1003: START - Cookbook sre.hosts.reimage for host cp5026.eqsin.wmnet with OS trixie
* 13:34 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 13:33 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5026.mgmt.eqsin.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:29 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp6002.mgmt.drmrs.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:25 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on dse-k8s-etcd[1001-1003].eqiad.wmnet with reason: Maintenance to increase vCPUS [[phab:T438084|T438084]]
* 13:24 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:24 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:22 slyngshede@cumin1003: START - Cookbook sre.hosts.provision for host cp5026.mgmt.eqsin.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342231{{!}}SI: Unset all filters on links to cases (T434530)]], [[gerrit:1342234{{!}}SI: Unset all filters on links to cases (T434530)]] (duration: 13m 10s)
* 13:19 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org
* 13:19 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org
* 13:19 sukhe@cumin1004: START - Cookbook sre.hosts.provision for host cp6002.mgmt.drmrs.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:18 stran@deploy1003: stran: Continuing with deployment
* 13:15 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: apply
* 13:15 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: apply
* 13:13 stran@deploy1003: stran: Backport for [[gerrit:1342231{{!}}SI: Unset all filters on links to cases (T434530)]], [[gerrit:1342234{{!}}SI: Unset all filters on links to cases (T434530)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:11 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp5026.eqsin.wmnet
* 13:09 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1342231{{!}}SI: Unset all filters on links to cases (T434530)]], [[gerrit:1342234{{!}}SI: Unset all filters on links to cases (T434530)]]
* 13:09 sukhe@cumin1004: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp6002.drmrs.wmnet
* 13:05 sukhe@cumin1004: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp6002.drmrs.wmnet
* 13:05 sukhe@cumin1004: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts cp6002.drmrs.wmnet
* 13:00 dkertesz@cumin1004: conftool action : set/pooled=yes; selector: name=cp5025.eqsin.wmnet
* 12:59 dkertesz@cumin1004: conftool action : set/weight=1; selector: name=cp5025.eqsin.wmnet
* 12:52 sukhe@cumin1004: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp6002.drmrs.wmnet
* 12:52 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cp6002.mgmt.drmrs.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:51 dkertesz: eqsin pooled again ([[phab:T438052|T438052]])
* 12:49 dkertesz@cumin1004: conftool action : set/pooled=yes; selector: cluster=dnsbox,dc=eqsin
* 12:47 dkertesz@dns1004: END - running authdns-update
* 12:45 dkertesz@dns1004: START - running authdns-update
* 12:43 dkertesz@cumin1004: conftool action : set/pooled=yes; selector: cluster=dnsbox,dc=eqsin,service=authdns-update
* 12:41 dkertesz@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqsin [reason: no reason specified, [[phab:T438052|T438052]]]
* 12:41 dkertesz@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool eqsin [reason: no reason specified, [[phab:T438052|T438052]]]
* 12:38 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a1-eqiad
* 12:38 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a1-eqiad
* 12:34 sukhe@cumin1004: START - Cookbook sre.hosts.provision for host cp6002.mgmt.drmrs.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:34 sukhe@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cp6002.drmrs.wmnet with reason: reimage
* 12:33 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp6002.drmrs.wmnet
* 12:13 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5025.eqsin.wmnet with OS trixie
* 12:12 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 12:11 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-a1-eqiad
* 12:09 cmooney@cumin1004: START - Cookbook sre.network.tls for network device ssw1-a1-eqiad
* 12:01 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.20 refs [[phab:T430839|T430839]]
* 11:59 moritzm: pruned obsolete Bullseye image python3-bullseye from the docker registry [[phab:T416452|T416452]]
* 11:50 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341284{{!}}IS/IS-labs: Set wmgUseModeratorToolkit default false (T431000)]] (duration: 10m 32s)
* 11:46 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ml-lab1002.eqiad.wmnet
* 11:45 samtar@deploy1003: samtar: Continuing with deployment
* 11:44 samtar@deploy1003: samtar: Backport for [[gerrit:1341284{{!}}IS/IS-labs: Set wmgUseModeratorToolkit default false (T431000)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:41 klausman@cumin1003: START - Cookbook sre.hosts.reboot-single for host ml-lab1002.eqiad.wmnet
* 11:39 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1341284{{!}}IS/IS-labs: Set wmgUseModeratorToolkit default false (T431000)]]
* 11:39 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5025.eqsin.wmnet with reason: host reimage
* 11:35 slyngshede@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5025.eqsin.wmnet with reason: host reimage
* 11:34 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:33 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:31 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:31 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:27 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:27 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:26 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:25 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:24 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.20 refs [[phab:T430839|T430839]]
* 11:21 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply
* 11:21 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply
* 11:20 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply
* 11:20 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply
* 11:19 moritzm: kicked off a new run of production-images-weekly-rebuild.service on build2004 (previously some leftovers of buster in the config prevented a complete run)
* 11:17 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
* 11:16 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
* 11:11 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
* 11:11 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
* 11:10 jnuche@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342210{{!}}Revert "Tell VisualEditor about the app web edit tags" (T437736 T438125)]] (duration: 33m 14s)
* 11:10 slyngshede@cumin1003: START - Cookbook sre.hosts.reimage for host cp5025.eqsin.wmnet with OS trixie
* 11:05 marostegui@cumin1004: dbctl commit (dc=all): 'Remove db1180 from dbctl [[phab:T437222|T437222]]', diff saved to https://phabricator.wikimedia.org/P96459 and previous config saved to /var/cache/conftool/dbconfig/20260916-110502-marostegui.json
* 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 11:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:01 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 11:01 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 10:57 jnuche@deploy1003: jnuche: Continuing with deployment
* 10:57 jnuche@deploy1003: jnuche: Backport for [[gerrit:1342210{{!}}Revert "Tell VisualEditor about the app web edit tags" (T437736 T438125)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:37 jnuche@deploy1003: Started scap sync-world: Backport for [[gerrit:1342210{{!}}Revert "Tell VisualEditor about the app web edit tags" (T437736 T438125)]]
* 10:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5025.mgmt.eqsin.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:10 slyngshede@cumin1003: START - Cookbook sre.hosts.provision for host cp5025.mgmt.eqsin.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:02 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 10:02 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 09:58 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 09:49 slyngshede@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on cp5025.eqsin.wmnet with reason: reimaging
* 09:48 slyngshede@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cp5025.eqsin.wmnet with reason: reimaging
* 09:41 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-timeline: apply
* 09:41 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-timeline: apply
* 09:38 moritzm: imported routinator 0.15.2-1trixie to thirdparty/routinator [[phab:T438122|T438122]]
* 09:30 slyngshede@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cp5025.mgmt.eqsin.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:29 slyngshede@cumin1003: START - Cookbook sre.hosts.provision for host cp5025.mgmt.eqsin.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:19 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 09:19 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp5025.eqsin.wmnet
* 09:19 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'sync'.
* 09:18 slyngshede@cumin1003: conftool action : set/pooled=yes; selector: name=cp3074.esams.wmnet
* 09:18 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp3074.esams.wmnet
* 09:15 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 09:12 elukey@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'sync'.
* 09:12 elukey@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'sync'.
* 09:11 elukey@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'sync'.
* 09:11 elukey@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'sync'.
* 09:10 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 09:10 elukey@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 09:09 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 09:09 elukey@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 08:55 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 08:54 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 08:40 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:40 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:36 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqsin [reason: depooling for maintainance, [[phab:T438052|T438052]]]
* 08:36 slyngshede@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool eqsin [reason: depooling for maintainance, [[phab:T438052|T438052]]]
* 08:35 slyngshede@cumin1003: END (FAIL) - Cookbook sre.dns.admin (exit_code=99) DNS admin: depool eqsin [reason: no reason specified, no task ID specified]
* 08:35 slyngshede@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool eqsin [reason: no reason specified, no task ID specified]
* 08:35 slyngshede@cumin1003: conftool action : set/pooled=no; selector: cluster=dnsbox,dc=eqsin
* 08:34 fabfur: start depooling eqsin ([[phab:T438052|T438052]])
* 08:24 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.20 refs [[phab:T430839|T430839]]
* 08:22 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:22 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:14 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.20 refs [[phab:T430839|T430839]]
* 08:11 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 08:11 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 08:11 jnuche@deploy1003: Rolling back deployment
* 08:10 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 08:07 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 07:59 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 07:59 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 07:58 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 07:54 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 07:34 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply
* 07:34 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply
* 07:20 mlitn@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342112{{!}}Adds an instrument for pre-image-carousel-retest (T437076)]], [[gerrit:1342113{{!}}Adds an instrument for pre-image-carousel-retest (T437076)]], [[gerrit:1342117{{!}}Set up instrument for 5-arm test (T437076)]], [[gerrit:1342118{{!}}Set up instrument for 5-arm test (T437076)]] (duration: 10m 56s)
* 07:16 mlitn@deploy1003: mlitn: Continuing with deployment
* 07:15 mlitn@deploy1003: mlitn: Backport for [[gerrit:1342112{{!}}Adds an instrument for pre-image-carousel-retest (T437076)]], [[gerrit:1342113{{!}}Adds an instrument for pre-image-carousel-retest (T437076)]], [[gerrit:1342117{{!}}Set up instrument for 5-arm test (T437076)]], [[gerrit:1342118{{!}}Set up instrument for 5-arm test (T437076)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be veri
* 07:09 mlitn@deploy1003: Started scap sync-world: Backport for [[gerrit:1342112{{!}}Adds an instrument for pre-image-carousel-retest (T437076)]], [[gerrit:1342113{{!}}Adds an instrument for pre-image-carousel-retest (T437076)]], [[gerrit:1342117{{!}}Set up instrument for 5-arm test (T437076)]], [[gerrit:1342118{{!}}Set up instrument for 5-arm test (T437076)]]
* 06:50 moritzm: installing sudo security updates
* 06:47 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply
* 06:37 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply
* 05:12 moritzm: pruned obsolete Bullseye image buildkitd from the docker registry [[phab:T416452|T416452]]
* 04:56 kart_: Updated Apertium to 2026-09-15-084320-production ([[phab:T437213|T437213]])
* 04:54 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/apertium: apply
* 04:54 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/apertium: apply
* 04:50 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/apertium: apply
* 04:49 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/apertium: apply
* 04:45 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/apertium: apply
* 04:45 kartik@deploy1003: helmfile [staging] START helmfile.d/services/apertium: apply
* 04:24 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 04:24 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 04:22 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 04:22 oblivian@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 36s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-15 ==
* 23:09 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply
* 23:08 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply
* 23:08 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply
* 23:08 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply
* 23:07 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 23:07 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 23:07 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 23:07 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 23:06 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 23:06 rzl@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 22:57 sukhe@puppetserver1001: conftool action : set/weight=1; selector: name=cp6001.drmrs.wmnet,service=cdn
* 22:50 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc-wf1002.eqiad.wmnet with OS trixie
* 22:46 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs ([[phab:T436363|T436363]])
* 22:46 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs6003.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:46 brett@cumin2003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs6003.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:46 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs6003.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:46 brett@cumin2003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs6003.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:45 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs6002.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:45 brett@cumin2003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs6002.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:45 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs6002.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:44 brett@cumin2003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs6002.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:44 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs6001.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:44 brett@cumin2003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs6001.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:43 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs6001.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:43 brett@cumin2003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs6001.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:43 brett@cumin2003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs ([[phab:T436363|T436363]])
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc-wf1002.eqiad.wmnet with reason: host reimage
* 22:33 brett@puppetserver1001: conftool action : set/weight=100; selector: name=cp6001.*
* 22:32 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp6001.*
* 22:26 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc-wf1002.eqiad.wmnet with reason: host reimage
* 22:07 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc-wf1002
* 22:07 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-wf1002
* 22:07 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host mc-wf1002
* 22:07 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc-wf1002.eqiad.wmnet 142.48.64.10.in-addr.arpa 2.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:07 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache mc-wf1002.eqiad.wmnet 142.48.64.10.in-addr.arpa 2.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:07 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:07 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc-wf1002 - rzl@cumin2003"
* 22:07 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc-wf1002 - rzl@cumin2003"
* 22:02 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:01 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host mc-wf1002
* 22:01 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host mc-wf1002.eqiad.wmnet with OS trixie
* 21:57 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6001.drmrs.wmnet with OS trixie
* 21:55 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply
* 21:55 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply
* 21:53 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply
* 21:53 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply
* 21:53 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:53 rzl@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:52 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:52 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:48 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:48 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:34 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6001.drmrs.wmnet with reason: host reimage
* 21:30 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp6001.drmrs.wmnet with reason: host reimage
* 21:20 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342049{{!}}MobileFrontend: Add app icons (T434258)]] (duration: 11m 47s)
* 21:15 jdlrobson@deploy1003: jdlrobson, cklimas: Continuing with deployment
* 21:13 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp6001.drmrs.wmnet with OS trixie
* 21:12 jdlrobson@deploy1003: jdlrobson, cklimas: Backport for [[gerrit:1342049{{!}}MobileFrontend: Add app icons (T434258)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:12 brett@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp6001.mgmt.drmrs.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:08 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1342049{{!}}MobileFrontend: Add app icons (T434258)]]
* 20:51 cdobbins@puppetserver1001: conftool action : set/weight=1; selector: name=cp2046.codfw.wmnet
* 20:51 cdobbins@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp2046.codfw.wmnet
* 20:48 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341271{{!}}Parsoid Read Views: Enable on all namespaces on wikitech (labswiki) (T437916)]] (duration: 09m 11s)
* 20:47 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp2046.codfw.wmnet with OS trixie
* 20:43 arlolra@deploy1003: ssastry, arlolra: Continuing with deployment
* 20:42 arlolra@deploy1003: ssastry, arlolra: Backport for [[gerrit:1341271{{!}}Parsoid Read Views: Enable on all namespaces on wikitech (labswiki) (T437916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:38 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1341271{{!}}Parsoid Read Views: Enable on all namespaces on wikitech (labswiki) (T437916)]]
* 20:34 brett@cumin2003: START - Cookbook sre.hosts.provision for host cp6001.mgmt.drmrs.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:30 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp6001.*
* 20:24 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp2046.codfw.wmnet with reason: host reimage
* 20:23 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339741{{!}}Enable ReaderExperiments in eswiki, jawiki, and ptwiki (T438009)]] (duration: 15m 58s)
* 20:20 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cp2046.codfw.wmnet with reason: host reimage
* 20:19 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T436363|T436363]])
* 20:19 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs4010.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:19 brett@cumin2003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs4010.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:19 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs4010.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:19 brett@cumin2003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs4010.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:18 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs4009.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:18 arlolra@deploy1003: lwatson, arlolra: Continuing with deployment
* 20:18 brett@cumin2003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs4009.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:17 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs4009.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:17 brett@cumin2003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs4009.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:17 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs4008.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:17 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp3075.esams.wmnet
* 20:17 sukhe@puppetserver1001: conftool action : set/weight=1; selector: name=cp3075.esams.wmnet
* 20:17 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp1103.eqiad.wmnet
* 20:17 brett@cumin2003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs4008.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:16 sukhe@puppetserver1001: conftool action : set/weight=1; selector: name=cp1103.eqiad.wmnet
* 20:16 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs4008.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:16 brett@cumin2003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs4008.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:16 brett@cumin2003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T436363|T436363]])
* 20:11 arlolra@deploy1003: lwatson, arlolra: Backport for [[gerrit:1339741{{!}}Enable ReaderExperiments in eswiki, jawiki, and ptwiki (T438009)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:10 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) config_reloading A:liberica-ulsfo ([[phab:T436363|T436363]])
* 20:08 brett@cumin2003: START - Cookbook sre.loadbalancer.admin config_reloading A:liberica-ulsfo ([[phab:T436363|T436363]])
* 20:08 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp1103.eqiad.wmnet with OS trixie
* 20:07 inflatador: bking@ganeti1046 sudo gnt-instance modify -B memory=4g,vcpus=4 dse-k8s-etcd100[1-3].eqiad.wmnet [[phab:T438084|T438084]]
* 20:07 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1339741{{!}}Enable ReaderExperiments in eswiki, jawiki, and ptwiki (T438009)]]
* 20:06 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp3075.esams.wmnet with OS trixie
* 20:04 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp7009.*
* 20:04 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host cp2046.codfw.wmnet with OS trixie
* 20:02 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp7009.*
* 20:02 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp7009.*
* 20:02 brett@puppetserver1001: conftool action : set/weight=1; selector: name=cp7009.*
* 20:01 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp7009.magru.wmnet with OS trixie
* 19:53 brett@puppetserver1001: conftool action : set/weight=1; selector: name=cp4045.*
* 19:53 brett@puppetserver1001: conftool action : set/weight=1; selector: name=cp4046.*
* 19:52 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4046.*
* 19:51 cdobbins@puppetserver1001: conftool action : set/pooled=no; selector: name=cp2046.codfw.wmnet
* 19:51 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp4046.ulsfo.wmnet with OS trixie
* 19:49 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4045.*
* 19:45 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp1103.eqiad.wmnet with reason: host reimage
* 19:43 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp4045.ulsfo.wmnet with OS trixie
* 19:41 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp3075.esams.wmnet with reason: host reimage
* 19:39 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cp1103.eqiad.wmnet with reason: host reimage
* 19:37 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp7009.magru.wmnet with reason: host reimage
* 19:33 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cp3075.esams.wmnet with reason: host reimage
* 19:32 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp7009.magru.wmnet with reason: host reimage
* 19:27 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp4046.ulsfo.wmnet with reason: host reimage
* 19:23 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp4046.ulsfo.wmnet with reason: host reimage
* 19:21 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp1103.eqiad.wmnet with OS trixie
* 19:19 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp1103.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:19 sukhe@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cp1103.eqiad.wmnet with reason: reimage
* 19:18 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp4045.ulsfo.wmnet with reason: host reimage
* 19:13 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 19:12 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 19:12 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 19:12 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp3075.esams.wmnet with OS trixie
* 19:12 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp4045.ulsfo.wmnet with reason: host reimage
* 19:12 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 19:11 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp3075.mgmt.esams.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:10 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp7009.magru.wmnet with OS trixie
* 19:09 brett@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp7009.mgmt.magru.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:08 sukhe@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cp3075.esams.wmnet with reason: reimaging
* 19:08 sukhe@cumin1004: START - Cookbook sre.hosts.provision for host cp1103.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:05 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp4046.ulsfo.wmnet with OS trixie
* 19:05 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore1006.eqiad.wmnet with OS bookworm
* 19:04 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3075.*
* 19:02 brett@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp4046.mgmt.ulsfo.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:01 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp1103.eqiad.wmnet
* 19:01 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp1103.eqiad.wmnet
* 19:00 sukhe@cumin1004: START - Cookbook sre.hosts.provision for host cp3075.mgmt.esams.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:58 brett@cumin2003: START - Cookbook sre.hosts.provision for host cp7009.mgmt.magru.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:57 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3075.esams.wmnet
* 18:55 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp1101.eqiad.wmnet
* 18:55 sukhe@puppetserver1001: conftool action : set/weight=1; selector: name=cp1101.eqiad.wmnet
* 18:55 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp4045.ulsfo.wmnet with OS trixie
* 18:52 brett@cumin2003: START - Cookbook sre.hosts.provision for host cp4046.mgmt.ulsfo.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:52 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp1101.eqiad.wmnet with OS trixie
* 18:45 cdobbins@puppetserver1001: conftool action : set/weight=1; selector: name=cp2044.codfw.wmnet
* 18:44 cdobbins@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp2044.codfw.wmnet
* 18:44 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore1006.eqiad.wmnet with reason: host reimage
* 18:40 brett@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp4045.mgmt.ulsfo.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:40 eevans@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore1006.eqiad.wmnet with reason: host reimage
* 18:39 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp3074.esams.wmnet
* 18:36 sukhe@cumin1004: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) config_reloading P<nowiki>{</nowiki>lvs3008.esams.wmnet<nowiki>}</nowiki> and A:liberica
* 18:36 sukhe@cumin1004: START - Cookbook sre.loadbalancer.admin config_reloading P<nowiki>{</nowiki>lvs3008.esams.wmnet<nowiki>}</nowiki> and A:liberica
* 18:33 sukhe@puppetserver1001: conftool action : set/weight=1; selector: name=cp3074.esams.wmnet
* 18:32 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp2044.codfw.wmnet with OS trixie
* 18:31 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp3074.esams.wmnet with OS trixie
* 18:30 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp1101.eqiad.wmnet with reason: host reimage
* 18:29 brett@cumin2003: START - Cookbook sre.hosts.provision for host cp4045.mgmt.ulsfo.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:26 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cp1101.eqiad.wmnet with reason: host reimage
* 18:22 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp7009.magru.wmnet with OS trixie
* 18:20 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1006.eqiad.wmnet with OS bookworm
* 18:20 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:19 eevans@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:19 eevans@cumin1004: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore1006.eqiad.wmnet
* 18:19 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore1006.eqiad.wmnet
* 18:10 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp1101.eqiad.wmnet with OS trixie
* 18:09 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp2044.codfw.wmnet with reason: host reimage
* 18:09 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp4045.ulsfo.wmnet with OS trixie
* 18:08 eevans@cumin1004: START - Cookbook sre.hosts.reboot-single for host sessionstore1006.eqiad.wmnet
* 18:07 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp3074.esams.wmnet with reason: host reimage
* 17:52 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp7009.magru.wmnet with reason: host reimage
* 17:48 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp2044.codfw.wmnet with OS trixie
* 17:43 cdobbins@puppetserver1001: conftool action : set/pooled=no; selector: name=cp2044.codfw.wmnet
* 17:36 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp3074.esams.wmnet with OS trixie
* 17:34 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp4046.ulsfo.wmnet with OS trixie
* 17:34 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp4045.ulsfo.wmnet with OS trixie
* 17:33 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3074.esams.wmnet
* 17:28 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp4046.*
* 17:28 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp4045.*
* 17:26 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp7009.magru.wmnet with OS trixie
* 17:25 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp1101.eqiad.wmnet with OS trixie
* 17:24 cdobbins@puppetserver1001: conftool action : set/pooled=no; selector: name=cp7009.magru.wmnet
* 17:23 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp1101.eqiad.wmnet
* 17:22 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:22 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org
* 17:02 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp7009.*
* 17:01 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:01 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:00 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:59 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:59 vriley@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host ms-be1099
* 16:59 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ms-be1099
* 16:59 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:56 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 16:55 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:55 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:55 vriley@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host ms-be1099
* 16:55 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ms-be1099
* 16:55 vriley@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host ms-be1099
* 16:54 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ms-be1099
* 16:54 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1098.eqiad.wmnet with OS bullseye
* 16:53 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:53 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [ms-be1099] - vriley@cumin1003"
* 16:53 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [ms-be1099] - vriley@cumin1003"
* 16:49 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 16:33 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS bullseye
* 16:21 mutante: temp disabling puppet on C:zookeeper (32 hosts) - safe deploy of https://gerrit.wikimedia.org/r/c/operations/puppet/+/1327569
* 16:04 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341890{{!}}Restore table borders for client-side MathJax (T435274)]], [[gerrit:1340558{{!}}lift IP cap for edit-a-thon /workshop (T437609 T437594 T437470)]] (duration: 24m 19s)
* 15:59 krinkle@deploy1003: anzx, krinkle: Continuing with deployment
* 15:44 krinkle@deploy1003: anzx, krinkle: Backport for [[gerrit:1341890{{!}}Restore table borders for client-side MathJax (T435274)]], [[gerrit:1340558{{!}}lift IP cap for edit-a-thon /workshop (T437609 T437594 T437470)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:40 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1341890{{!}}Restore table borders for client-side MathJax (T435274)]], [[gerrit:1340558{{!}}lift IP cap for edit-a-thon /workshop (T437609 T437594 T437470)]]
* 15:34 elukey@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 15:34 elukey@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 15:33 brennen@deploy1003: Finished deploy [phabricator/deployment@c386249]: deploy phab1005 for [[phab:T437930|T437930]] (duration: 00m 39s)
* 15:33 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS bullseye
* 15:33 brennen@deploy1003: Started deploy [phabricator/deployment@c386249]: deploy phab1005 for [[phab:T437930|T437930]]
* 15:32 elukey@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 15:32 elukey@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'sync'.
* 15:32 brennen@deploy1003: Finished deploy [phabricator/deployment@c386249]: deploy phab2003 for [[phab:T437930|T437930]] (duration: 00m 52s)
* 15:32 elukey@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 15:32 elukey@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'sync'.
* 15:31 brennen@deploy1003: Started deploy [phabricator/deployment@c386249]: deploy phab2003 for [[phab:T437930|T437930]]
* 15:31 elukey@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'.
* 15:31 elukey@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'.
* 15:26 jelto@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on phab2003.codfw.wmnet,phab[1005-1006].eqiad.wmnet with reason: Phabricator deploy
* 15:26 moritzm: pruned obsolete Bullseye image amd-gpu-tester from the docker registry [[phab:T416452|T416452]]
* 15:12 elukey@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 15:12 elukey@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 15:11 elukey@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 15:11 elukey@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 15:00 tgr@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341292{{!}}CommonSettings: Use a restrictive CSP for auth.wikimedia.org (T419684)]] (duration: 25m 11s)
* 14:55 tgr@deploy1003: tgr, arendpieter: Continuing with deployment
* 14:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:53 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:53 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:52 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:49 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 14:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 14:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 14:47 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 14:47 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 14:45 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 14:45 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 14:45 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 14:44 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 14:42 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 14:42 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 14:39 tgr@deploy1003: tgr, arendpieter: Backport for [[gerrit:1341292{{!}}CommonSettings: Use a restrictive CSP for auth.wikimedia.org (T419684)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:34 tgr@deploy1003: Started scap sync-world: Backport for [[gerrit:1341292{{!}}CommonSettings: Use a restrictive CSP for auth.wikimedia.org (T419684)]]
* 14:17 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341895{{!}}ReportIncidentController: Instance cache expensive methods (T437588)]] (duration: 11m 56s)
* 14:16 btullis@cumin1004: END (PASS) - Cookbook sre.ceph.rotate-osd-keys (exit_code=0) rolling rotate_keys on A:cephosd-codfw
* 14:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:12 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 14:09 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1341895{{!}}ReportIncidentController: Instance cache expensive methods (T437588)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:05 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1341895{{!}}ReportIncidentController: Instance cache expensive methods (T437588)]]
* 13:43 btullis@cumin1004: START - Cookbook sre.ceph.rotate-osd-keys rolling rotate_keys on A:cephosd-codfw
* 13:36 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341861{{!}}SuggestedInvestigations: Update "sockpuppet" queue view defaults (T438018)]] (duration: 10m 23s)
* 13:32 stran@deploy1003: stran: Continuing with deployment
* 13:30 stran@deploy1003: stran: Backport for [[gerrit:1341861{{!}}SuggestedInvestigations: Update "sockpuppet" queue view defaults (T438018)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:26 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1341861{{!}}SuggestedInvestigations: Update "sockpuppet" queue view defaults (T438018)]]
* 13:21 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:20 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:19 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334946{{!}}ArticleGuidance: Remove the experiment configuration keys (T434487)]] (duration: 09m 19s)
* 13:16 btullis@cumin1004: END (PASS) - Cookbook sre.ceph.rotate-osd-keys (exit_code=0) rolling rotate_keys on P<nowiki>{</nowiki>cephosd2001.codfw.wmnet<nowiki>}</nowiki> and (A:cephosd-codfw or A:cephosd-eqiad)
* 13:15 sbisson@deploy1003: sbisson: Continuing with deployment
* 13:14 sbisson@deploy1003: sbisson: Backport for [[gerrit:1334946{{!}}ArticleGuidance: Remove the experiment configuration keys (T434487)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) config_reloading P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:10 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1334946{{!}}ArticleGuidance: Remove the experiment configuration keys (T434487)]]
* 13:10 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin config_reloading P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:08 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:08 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:07 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:07 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:07 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:07 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:03 btullis@cumin1004: START - Cookbook sre.ceph.rotate-osd-keys rolling rotate_keys on P<nowiki>{</nowiki>cephosd2001.codfw.wmnet<nowiki>}</nowiki> and (A:cephosd-codfw or A:cephosd-eqiad)
* 13:00 btullis@cumin1004: END (PASS) - Cookbook sre.ceph.rotate-osd-keys (exit_code=0) rolling rotate_keys on P<nowiki>{</nowiki>cephosd2001.codfw.wmnet<nowiki>}</nowiki> and (A:cephosd-codfw or A:cephosd-eqiad)
* 12:59 btullis@cumin1004: START - Cookbook sre.ceph.rotate-osd-keys rolling rotate_keys on P<nowiki>{</nowiki>cephosd2001.codfw.wmnet<nowiki>}</nowiki> and (A:cephosd-codfw or A:cephosd-eqiad)
* 12:46 btullis@cumin1004: END (PASS) - Cookbook sre.ceph.rotate-osd-keys (exit_code=0) rolling rotate_keys on P<nowiki>{</nowiki>cephosd2001.codfw.wmnet<nowiki>}</nowiki> and (A:cephosd-codfw or A:cephosd-eqiad)
* 12:45 btullis@cumin1004: START - Cookbook sre.ceph.rotate-osd-keys rolling rotate_keys on P<nowiki>{</nowiki>cephosd2001.codfw.wmnet<nowiki>}</nowiki> and (A:cephosd-codfw or A:cephosd-eqiad)
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: network maintenance finished, [[phab:T437984|T437984]]]
* 12:40 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: network maintenance finished, [[phab:T437984|T437984]]]
* 12:29 XioNoX: asw1-b4-magru> request system reboot - [[phab:T437984|T437984]]
* 12:24 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 12:24 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 12:24 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 12:23 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 12:23 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 12:23 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 12:22 moritzm: installing shadow security updates
* 12:19 slyngshede@puppetserver1001: conftool action : set/weight=1; selector: name=cp7010.magru.wmnet
* 12:13 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 12:13 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 12 hosts with reason: Switch maintenance
* 12:12 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-b4-magru,asw1-b4-magru IPv6,asw1-b4-magru.mgmt with reason: Switch maintenance
* 12:11 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 12:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: switch reboot, [[phab:T437984|T437984]]]
* 12:11 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: switch reboot, [[phab:T437984|T437984]]]
* 12:09 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on install7002.wikimedia.org with reason: switch reboot
* 12:08 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 12:07 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 12:07 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 12:07 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 12:06 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 12:06 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 12:03 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 12:03 XioNoX: push pfw policies - [[phab:T437627|T437627]]
* 12:01 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 12:01 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) config_reloading P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 12:00 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin config_reloading P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 11:56 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 11:56 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 11:33 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2250.codfw.wmnet with reason: cloning db2201
* 11:18 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7004.magru.wmnet
* 11:17 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7004.magru.wmnet
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 11:05 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 11:05 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 11:04 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 10:51 slyngshede@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp7010.magru.wmnet
* 10:34 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 10:33 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 10:30 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1098.eqiad.wmnet with OS trixie
* 10:21 moritzm: failover Ganeti master in magru to ganeti7001
* 10:20 moritzm: increased DRBD replication speed in Ganeti/magru [[phab:T428878|T428878]]
* 10:10 hashar@deploy1003: Finished deploy [integration/docroot@5cf09c8]: build: Updating npm dependencies (duration: 00m 13s)
* 10:10 hashar@deploy1003: Started deploy [integration/docroot@5cf09c8]: build: Updating npm dependencies
* 10:09 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 10:08 moritzm: increased DRBD replication speed in Ganeti/esams [[phab:T428878|T428878]]
* 10:07 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 10:05 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 10:05 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:39 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: switches reboot, [[phab:T437984|T437984]]]
* 09:39 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool esams [reason: switches reboot, [[phab:T437984|T437984]]]
* 09:31 XioNoX: asw1-by27-esams> request system reboot - [[phab:T437984|T437984]]
* 09:30 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS trixie
* 09:28 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp7010.magru.wmnet with OS trixie
* 09:26 ayounsi@cumin1003: END (FAIL) - Cookbook sre.network.depool-rack (exit_code=99) with action 'depool' for esams rack BY27
* 09:24 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for esams rack BY27
* 09:24 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1098.eqiad.wmnet with OS trixie
* 09:23 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS trixie
* 09:22 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be1098.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 09:15 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.20 refs [[phab:T430839|T430839]]
* 09:10 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1098.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 09:06 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-by27-esams,asw1-by27-esams IPv6,asw1-by27-esams.mgmt with reason: Switch maintenance
* 09:05 ayounsi@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on asw1-by27-esams IPv6,asw1-by27-esams.mgmt,asw1-by-27-esams with reason: Switch maintenance
* 09:04 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 12 hosts with reason: Switch maintenance
* 09:04 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp7010.magru.wmnet with reason: host reimage
* 09:01 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: switches reboot, [[phab:T437984|T437984]]]
* 09:00 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: switches reboot, [[phab:T437984|T437984]]]
* 09:00 slyngshede@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp7010.magru.wmnet with reason: host reimage
* 08:59 jnuche@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341698{{!}}RestSandbox: Pass JsonLocalizer instead of ResponseFactory to ModuleManager (T437982)]] (duration: 12m 03s)
* 08:55 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2197.codfw.wmnet with reason: cloning db2201
* 08:55 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on install3004.wikimedia.org with reason: switch reboot
* 08:53 jnuche@deploy1003: jnuche: Continuing with deployment
* 08:52 jnuche@deploy1003: jnuche: Backport for [[gerrit:1341698{{!}}RestSandbox: Pass JsonLocalizer instead of ResponseFactory to ModuleManager (T437982)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:49 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: sync
* 08:49 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: sync
* 08:47 jnuche@deploy1003: Started scap sync-world: Backport for [[gerrit:1341698{{!}}RestSandbox: Pass JsonLocalizer instead of ResponseFactory to ModuleManager (T437982)]]
* 08:33 slyngshede@cumin1003: START - Cookbook sre.hosts.reimage for host cp7010.magru.wmnet with OS trixie
* 08:26 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be1098.eqiad.wmnet with OS trixie
* 08:26 slyngshede@puppetserver1001: conftool action : set/pooled=no; selector: name=cp7010.magru.wmnet
* 08:25 XioNoX: asw1-b3-magru - Disable logging and file logging for BRCM_PKT - [[phab:T437984|T437984]]
* 08:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS trixie
* 08:18 dpogorzelski@dns1004: END - running authdns-update
* 08:15 dpogorzelski@dns1004: START - running authdns-update
* 08:14 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3005.esams.wmnet
* 08:13 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3005.esams.wmnet
* 08:07 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341274{{!}}SI: Implement "queue view" functionality (T437183)]], [[gerrit:1341242{{!}}SuggestedInvestigations: Add and enable 'sockpuppets' queue view (T437183)]], [[gerrit:1341254{{!}}Add wmf-specific Special:SuggestedInvestigations messages (T437183)]] (duration: 55m 27s)
* 07:54 stran@deploy1003: stran: Continuing with deployment
* 07:31 stran@deploy1003: stran: Backport for [[gerrit:1341274{{!}}SI: Implement "queue view" functionality (T437183)]], [[gerrit:1341242{{!}}SuggestedInvestigations: Add and enable 'sockpuppets' queue view (T437183)]], [[gerrit:1341254{{!}}Add wmf-specific Special:SuggestedInvestigations messages (T437183)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:18 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply
* 07:18 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply
* 07:11 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1341274{{!}}SI: Implement "queue view" functionality (T437183)]], [[gerrit:1341242{{!}}SuggestedInvestigations: Add and enable 'sockpuppets' queue view (T437183)]], [[gerrit:1341254{{!}}Add wmf-specific Special:SuggestedInvestigations messages (T437183)]]
* 07:06 moritzm: pruned obsolete Bullseye image python3-devel from the docker registry [[phab:T416452|T416452]]
* 06:51 moritzm: pruned obsolete Bullseye image python3-build-bullseye from the docker registry [[phab:T416452|T416452]]
* 05:59 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply
* 05:49 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply
* 05:48 oblivian@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 05:47 oblivian@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 05:38 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply
* 05:28 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-video: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-video: apply
* 05:10 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-video: apply
* 05:10 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-video: apply
* 05:10 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-video: apply
* 05:10 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-video: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-syntaxhighlight: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-syntaxhighlight: apply
* 05:10 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-syntaxhighlight: apply
* 05:10 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-syntaxhighlight: apply
* 05:10 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-syntaxhighlight: apply
* 05:10 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-syntaxhighlight: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-media: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-media: apply
* 05:10 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-media: apply
* 05:10 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-media: apply
* 05:10 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-media: apply
* 05:10 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-media: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-constraints: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-constraints: apply
* 05:10 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-constraints: apply
* 05:09 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-constraints: apply
* 05:09 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-constraints: apply
* 05:09 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-constraints: apply
* 05:08 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply
* 05:08 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox: apply
* 05:07 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply
* 05:07 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox: apply
* 05:07 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox: apply
* 05:07 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox: apply
* 05:07 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply
* 05:07 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply
* 05:06 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-timeline: apply
* 05:06 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-timeline: apply
* 05:06 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply
* 05:06 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply
* 04:07 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.17 (duration: 07m 10s)
* 03:06 mwpresync@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.19,1.47.0-wmf.20,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.20 refs [[phab:T430839|T430839]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 22s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:43 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 00:23 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore1005.eqiad.wmnet with reason: host reimage
* 00:19 jclark@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore1005.eqiad.wmnet with reason: host reimage
* 00:17 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 00:07 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sessionstore1005.eqiad.wmnet with OS bookworm
== 2026-09-14 ==
* 23:41 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 23:26 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324966{{!}}Enable Produnto on pilot wikis (T421436)]] (duration: 12m 59s)
* 23:22 tstarling@deploy1003: tstarling: Continuing with deployment
* 23:17 tstarling@deploy1003: tstarling: Backport for [[gerrit:1324966{{!}}Enable Produnto on pilot wikis (T421436)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:13 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1324966{{!}}Enable Produnto on pilot wikis (T421436)]]
* 23:01 eevans@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 22:41 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341385{{!}}VisualEditor: don't register settings tool in wikitextCommandRegistry (T437810)]] (duration: 09m 22s)
* 22:36 kemayo@deploy1003: kemayo: Continuing with deployment
* 22:36 kemayo@deploy1003: kemayo: Backport for [[gerrit:1341385{{!}}VisualEditor: don't register settings tool in wikitextCommandRegistry (T437810)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:31 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1341385{{!}}VisualEditor: don't register settings tool in wikitextCommandRegistry (T437810)]]
* 22:22 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 22:22 eevans@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 21:43 sbassett: Deployed security fix for [[phab:T435623|T435623]]
* 21:29 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1098.eqiad.wmnet with OS bullseye
* 21:29 sbassett: Deployed security fix for [[phab:T434372|T434372]]
* 21:26 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 21:26 eevans@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 21:05 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338999{{!}}Enable ReadingLists for all logged-in users on English Wikipedia (T434923)]], [[gerrit:1340004{{!}}Enable Reading Recommendations experiment on test wiki (T437665)]] (duration: 11m 03s)
* 21:00 aude@deploy1003: aude, jdlrobson: Continuing with deployment
* 20:58 aude@deploy1003: aude, jdlrobson: Backport for [[gerrit:1338999{{!}}Enable ReadingLists for all logged-in users on English Wikipedia (T434923)]], [[gerrit:1340004{{!}}Enable Reading Recommendations experiment on test wiki (T437665)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:54 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1338999{{!}}Enable ReadingLists for all logged-in users on English Wikipedia (T434923)]], [[gerrit:1340004{{!}}Enable Reading Recommendations experiment on test wiki (T437665)]]
* 20:47 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341278{{!}}[A11y] Add list semantics to ReadingList page (T435864 T434923)]] (duration: 12m 49s)
* 20:43 aude@deploy1003: aude, jdlrobson: Continuing with deployment
* 20:39 aude@deploy1003: aude, jdlrobson: Backport for [[gerrit:1341278{{!}}[A11y] Add list semantics to ReadingList page (T435864 T434923)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:34 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1341278{{!}}[A11y] Add list semantics to ReadingList page (T435864 T434923)]]
* 20:32 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339811{{!}}wmf-config: Register content/v2-beta REST module as disabled (T432798)]], [[gerrit:1338274{{!}}wikifunctions: Move abstract fragments to mainstash (T432849)]] (duration: 25m 21s)
* 20:28 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:27 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:27 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:27 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:25 jforrester@deploy1003: jforrester, aghirelli: Continuing with deployment
* 20:24 jforrester@deploy1003: jforrester, aghirelli: Backport for [[gerrit:1339811{{!}}wmf-config: Register content/v2-beta REST module as disabled (T432798)]], [[gerrit:1338274{{!}}wikifunctions: Move abstract fragments to mainstash (T432849)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS bullseye
* 20:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be1098.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 20:06 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1339811{{!}}wmf-config: Register content/v2-beta REST module as disabled (T432798)]], [[gerrit:1338274{{!}}wikifunctions: Move abstract fragments to mainstash (T432849)]]
* 20:04 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1098.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 20:03 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be1098
* 20:02 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ms-be1098
* 20:02 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 20:02 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [ms-be1098] - vriley@cumin1003"
* 20:02 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [ms-be1098] - vriley@cumin1003"
* 19:59 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 19:58 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 19:57 dzahn@dns1005: END - running authdns-update
* 19:55 eevans@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 19:55 dzahn@dns1005: START - running authdns-update
* 19:54 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 19:38 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341273{{!}}[CodeMirror] enable for new users (enwiki), new VE integration (global) (T288161 T432558)]] (duration: 33m 51s)
* 19:26 musikanimal@deploy1003: musikanimal: Continuing with deployment
* 19:22 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1341273{{!}}[CodeMirror] enable for new users (enwiki), new VE integration (global) (T288161 T432558)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:04 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1341273{{!}}[CodeMirror] enable for new users (enwiki), new VE integration (global) (T288161 T432558)]]
* 18:53 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:50 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 18:50 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 18:49 eevans@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 18:26 brett@cumin2003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d6-eqiad
* 18:26 brett@cumin2003: START - Cookbook sre.network.tls for network device lsw1-d6-eqiad
* 18:26 brett@cumin2003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-eqiad
* 18:26 brett@cumin2003: START - Cookbook sre.network.tls for network device ssw1-d8-eqiad
* 18:25 brett@cumin2003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-eqiad
* 18:25 brett@cumin2003: START - Cookbook sre.network.tls for network device lsw1-d4-eqiad
* 18:25 root@cumin2003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-eqiad
* 18:25 root@cumin2003: START - Cookbook sre.network.tls for network device lsw1-d2-eqiad
* 18:17 jclark@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:14 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337611{{!}}extension-list: Add ModeratorToolkit (T431000)]] (duration: 09m 34s)
* 18:10 samtar@deploy1003: samtar: Continuing with deployment
* 18:09 samtar@deploy1003: samtar: Backport for [[gerrit:1337611{{!}}extension-list: Add ModeratorToolkit (T431000)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:06 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 18:05 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1337611{{!}}extension-list: Add ModeratorToolkit (T431000)]]
* 18:01 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 17:48 jclark@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 17:07 tgr@deploy1003: Finished scap sync-world: Backport for [[gerrit:1330446{{!}}CommonSettings: Use a restrictive, eval-free CSP for auth.wikimedia.org (T419684)]] (duration: 23m 19s)
* 17:00 tgr@deploy1003: arendpieter, tgr: Rolling back deployment
* 16:52 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 16:51 eevans@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 16:49 tgr@deploy1003: arendpieter, tgr: Backport for [[gerrit:1330446{{!}}CommonSettings: Use a restrictive, eval-free CSP for auth.wikimedia.org (T419684)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:44 tgr@deploy1003: Started scap sync-world: Backport for [[gerrit:1330446{{!}}CommonSettings: Use a restrictive, eval-free CSP for auth.wikimedia.org (T419684)]]
* 16:09 Amir1: drop links tables from db1252 ([[phab:T437278|T437278]])
* 16:07 Amir1: drop links tables from db2240 ([[phab:T437278|T437278]])
* 16:05 Amir1: drop non-links tables from db2247 ([[phab:T437278|T437278]])
* 15:53 Lucas_WMDE: UTC afternoon backport+config window belatedly done
* 15:50 lucaswerkmeister-wmde@deploy1003: mwscript-k8s job started: namespaceDupes abstractwiki --fix # [[phab:T437772|T437772]]
* 15:49 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335401{{!}}Adjust extendedconfirmed calculation to first edit on viwiki (T437006)]], [[gerrit:1340505{{!}}core-Namespaces: Add AW and AWT alias for its talk in abstractwiki (T437772)]] (duration: 10m 23s)
* 15:48 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 15:47 eevans@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 15:45 lucaswerkmeister-wmde@deploy1003: bunnypranav, lucaswerkmeister-wmde, tryvix1509: Continuing with deployment
* 15:43 lucaswerkmeister-wmde@deploy1003: bunnypranav, lucaswerkmeister-wmde, tryvix1509: Backport for [[gerrit:1335401{{!}}Adjust extendedconfirmed calculation to first edit on viwiki (T437006)]], [[gerrit:1340505{{!}}core-Namespaces: Add AW and AWT alias for its talk in abstractwiki (T437772)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:39 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1335401{{!}}Adjust extendedconfirmed calculation to first edit on viwiki (T437006)]], [[gerrit:1340505{{!}}core-Namespaces: Add AW and AWT alias for its talk in abstractwiki (T437772)]]
* 15:36 elukey@dns1004: END - running authdns-update
* 15:36 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 15:35 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 15:35 moritzm: installing shadow security updates
* 15:35 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 15:34 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:33 elukey@dns1004: START - running authdns-update
* 15:33 eevans@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:32 eevans@cumin1004: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore1005.eqiad.wmnet
* 15:29 lucaswerkmeister-wmde@deploy1003: mwscript-k8s job started: namespaceDupes afwiki --fix # [[phab:T437576|T437576]]
* 15:29 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338902{{!}}afwiki: Create Draft and Draft talk namespaces (T437576)]] (duration: 15m 44s)
* 15:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 15:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 15:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 15:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 15:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 15:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 15:21 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, tryvix1509: Continuing with deployment
* 15:21 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:17 eevans@cumin1004: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore1005.eqiad.wmnet
* 15:17 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, tryvix1509: Backport for [[gerrit:1338902{{!}}afwiki: Create Draft and Draft talk namespaces (T437576)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:17 eevans@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore1005.eqiad.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 15:13 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore1004.eqiad.wmnet
* 15:13 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1338902{{!}}afwiki: Create Draft and Draft talk namespaces (T437576)]]
* 15:06 eevans@cumin1004: START - Cookbook sre.hosts.reboot-single for host sessionstore1004.eqiad.wmnet
* 14:59 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore1004.eqiad.wmnet with OS bookworm
* 14:38 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore1004.eqiad.wmnet with reason: host reimage
* 14:33 marostegui@dns1004: END - running authdns-update
* 14:32 eevans@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore1004.eqiad.wmnet with reason: host reimage
* 14:30 marostegui@dns1004: START - running authdns-update
* 14:15 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1004.eqiad.wmnet with OS bookworm
* 14:14 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore1004.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:13 eevans@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1004.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:13 eevans@cumin1004: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore1004.eqiad.wmnet
* 14:13 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore1004.eqiad.wmnet
* 14:05 moritzm: kick off a rebuild of base images on build2004
* 14:05 moritzm: kick off a rebuild of base images on build2004
* 14:00 eevans@cumin1004: START - Cookbook sre.hosts.reboot-single for host sessionstore1004.eqiad.wmnet
* 14:00 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:00 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 13:56 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:54 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:52 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:44 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 13:43 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 13:42 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:40 eevans@cumin1004: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore1004.eqiad.wmnet
* 13:40 eevans@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore1004.eqiad.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 13:40 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:40 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:39 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:39 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:39 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 13:38 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 13:38 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:38 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:38 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 13:37 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 13:37 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:36 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:36 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 13:35 oblivian@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:35 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 13:35 oblivian@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:35 oblivian@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:34 oblivian@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'.
* 13:34 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 13:34 oblivian@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:34 oblivian@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:34 oblivian@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:34 oblivian@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 13:34 oblivian@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'.
* 13:33 oblivian@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'.
* 13:33 oblivian@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
* 13:33 oblivian@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
* 13:33 oblivian@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 13:33 oblivian@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 13:33 oblivian@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 13:32 oblivian@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 13:32 oblivian@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 13:32 oblivian@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 13:32 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 13:32 oblivian@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 13:32 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 13:32 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 13:29 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:23 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1003.eqiad.wmnet
* 13:21 sukhe: sudo cumin -b11 "A:cp-text" "run-puppet-agent --enable 'merging CR 1338134'" [[phab:T425441|T425441]]
* 13:20 sukhe: sudo cumin -b11 "A:cp-text" "run-puppet-agent --enable 'merging CR 1338134'"[[phab:T425441|T425441]]
* 13:19 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:17 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1003.eqiad.wmnet
* 13:14 moritzm: installing Bird security updates
* 13:09 sukhe: sudo cumin "A:cp-text" "disable-puppet 'merging CR 1338134'"
* 13:06 jmm@dns1004: END - running authdns-update
* 13:04 jmm@dns1004: START - running authdns-update
* 12:58 moritzm: update Trixie installer image to 13.7 [[phab:T437715|T437715]]
* 12:58 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 12:54 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 12:52 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 12:49 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 12:48 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 12:47 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 12:44 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 12:44 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 12:42 oblivian@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 12:40 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 12:40 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 12:40 oblivian@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 12:39 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 12:39 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 12:39 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 12:37 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 12:37 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 12:36 marostegui@cumin1004: conftool action : set/pooled=yes; selector: name=clouddb1025.eqiad.wmnet,service=x4
* 12:34 _joe_: adding gvisor labels to all wikikube clusters nodes
* 12:30 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 12:14 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 12:14 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 11:22 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1260: After cloning
* 10:48 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:45 ladsgroup@dns1004: END - running authdns-update
* 10:42 ladsgroup@dns1004: START - running authdns-update
* 10:37 marostegui@cumin1004: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=x4
* 10:37 marostegui@cumin1004: START - Cookbook sre.mysql.pool pool db1260: After cloning
* 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 10:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 10:04 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 10:04 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 09:53 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:53 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:41 Amir1: drop links tables from db2172 ([[phab:T437278|T437278]])
* 09:40 Amir1: drop links tables from db1228 ([[phab:T437278|T437278]])
* 09:08 marostegui: Stop mariadb on db1260 to clone dbstore1007, there will be lag on wikireplicas:x4 https://phabricator.wikimedia.org/T437839
* 09:07 taavi@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1025.eqiad.wmnet
* 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1260: Needs to clone another host from this one
* 09:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1260: Needs to clone another host from this one
* 09:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb[1024-1025].eqiad.wmnet,db[1155,1260].eqiad.wmnet,dbstore1007.eqiad.wmnet with reason: Adding x4
* 08:44 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:42 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:41 moritzm: pruned obsolete Bullseye images php8.3-icu72-cli / php8.3-icu72-fpm-multiversion-base / php8.3-icu72-fpm from the docker registry [[phab:T416452|T416452]]
* 08:37 taavi@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet
* 08:37 taavi@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1024.eqiad.wmnet
* 08:36 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on dbstore1007.eqiad.wmnet with reason: Adding x4
* 08:35 moritzm: pruned obsolete Bullseye images php8.1-cli/php8.1-fpm/ php8.1-fpm-multiversion-base from the docker registry [[phab:T416452|T416452]]
* 08:10 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: sync
* 08:08 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: sync
* 07:58 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 23s)
* 07:57 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 07:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 07:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 07:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 07:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 07:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 07:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 07:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 07:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 07:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 07:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 07:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 07:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 07:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 07:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 07:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 07:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 07:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 07:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 07:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 07:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 07:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 07:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 07:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 07:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 07:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 07:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 07:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 07:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 07:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 07:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 07:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 07:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 07:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 07:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 07:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 07:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 07:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 07:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 07:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 07:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 07:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 07:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 07:31 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1340801{{!}}Allow title-like strings with Package: prefix in require() (T430644)]], [[gerrit:1340802{{!}}Runtime: Add a facility for loading files by title (T430644)]] (duration: 34m 30s)
* 07:18 tstarling@deploy1003: tstarling: Continuing with deployment
* 07:17 tstarling@deploy1003: tstarling: Backport for [[gerrit:1340801{{!}}Allow title-like strings with Package: prefix in require() (T430644)]], [[gerrit:1340802{{!}}Runtime: Add a facility for loading files by title (T430644)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 06:56 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1340801{{!}}Allow title-like strings with Package: prefix in require() (T430644)]], [[gerrit:1340802{{!}}Runtime: Add a facility for loading files by title (T430644)]]
* 06:26 TimStarling: on deploy1003: docker image pull docker-registry.wikimedia.org/php8.3-fpm-multiversion-base
* 05:51 _joe_: pulled bookworm:latest from build2004 to build2001 [[phab:T437829|T437829]]
* 05:39 _joe_: force-running build-base-images on build2004 for [[phab:T437829|T437829]]
* 04:53 TimStarling: on build2001 rebuilding base images [[phab:T437829|T437829]]
* 03:00 tstarling@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.18,1.47.0-wmf.19,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted
* 02:59 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1340801{{!}}Allow title-like strings with Package: prefix in require() (T430644)]], [[gerrit:1340802{{!}}Runtime: Add a facility for loading files by title (T430644)]]
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-13 ==
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 29s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-12 ==
* 19:40 ladsgroup@cumin1003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-eqiad
* 19:32 ladsgroup@cumin1003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-eqiad
* 19:30 ladsgroup@cumin1003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw
* 19:21 ladsgroup@cumin1003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 35s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-11 ==
* 21:51 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 21:50 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 16:47 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339808{{!}}Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813{{!}}Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]] (duration: 07m 23s)
* 16:43 sbassett@deploy1003: sbassett: Continuing with deployment
* 16:42 sbassett@deploy1003: sbassett: Backport for [[gerrit:1339808{{!}}Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813{{!}}Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:40 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1339808{{!}}Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813{{!}}Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]]
* 16:08 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 14:39 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 14:10 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:10 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:10 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:09 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 13:40 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:28 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:11 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:10 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: db1199 repool
* 11:05 moritzm: installing Linux 6.1.187 on Bookworm hosts
* 11:05 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Jcrespo out of all services on: 2443 hosts
* 10:44 aokoth@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2020 in turn
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2019 in turn
* 10:41 Emperor: restart versitygw@objectstorage0[0-3].service on backup2018 in turn
* 10:40 Emperor: restart versitygw@objectstorage0[0-3].service on backup2017 in turn
* 10:39 Emperor: restart versitygw@objectstorage0[0-3].service on backup2016 in turn
* 10:37 Emperor: restart versitygw@objectstorage0[0-3].service on backup2015 in turn
* 10:36 Emperor: restart versitygw@objectstorage0[0-3].service on backup1020 in turn
* 10:35 Emperor: restart versitygw@objectstorage0[0-3].service on backup1019 in turn
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1199: db1199 repool
* 10:33 Emperor: restart versitygw@objectstorage0[0-3].service on backup1018 in turn
* 10:32 Emperor: restart versitygw@objectstorage0[0-3].service on backup1017 in turn
* 10:30 Emperor: restart versitygw@objectstorage0[0-3].service on backup1016 in turn
* 10:20 Emperor: restart versitygw@objectstorage0[1-3].service on backup1015 in turn
* 10:17 Emperor: restart versitygw@objectstorage00.service on backup1015
* 10:15 aokoth@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 08:46 slyngs: Update CAS/SSO to CAS 7.3.8.3
* 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:45 slyngshede@dns1004: END - running authdns-update
* 08:43 slyngshede@dns1004: START - running authdns-update
* 08:36 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:28 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 7 hosts with reason: Restarting s5
* 08:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db[1154,1269].eqiad.wmnet with reason: Restarting s5
* 08:20 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:20 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Repooling db1159
* 07:58 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:58 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:54 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:28 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:27 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1199: Needs to clone another host from this one
* 07:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1199: Needs to clone another host from this one
* 07:16 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Repooling db1159
* 07:15 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1199.eqiad.wmnet with reason: Cloning s4
* 07:10 TimStarling: killed jobs for [[phab:T437056|T437056]] since they weren't purging
* 06:38 TimStarling: also started refreshLinks for ptwiki and zhwiki, reparsing ~3000 pages altogether [[phab:T437056|T437056]]
* 06:27 TimStarling: for [[phab:T437056|T437056]]: mwscript-k8s refreshLinks.php --wiki=eswiki --tracking-category scribunto-common-error-category
* 05:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1159: Needs to clone another host from this one
* 05:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1159: Needs to clone another host from this one
* 05:30 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1159.eqiad.wmnet with reason: Cloning
* 05:29 marostegui: Start cloning db1245:s5 [[phab:T437563|T437563]]
* 05:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet,db1245.eqiad.wmnet with reason: Cloning
* 05:25 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] (duration: 09m 59s)
* 05:21 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:20 tstarling@deploy1003: tstarling: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:15 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 50s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-10 ==
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1349.eqiad.wmnet
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1349.eqiad.wmnet
* 23:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1349.eqiad.wmnet
* 23:09 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1349
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1349
* 22:32 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1349
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:28 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1349
* 22:27 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:27 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1349.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1348.eqiad.wmnet
* 22:23 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 22:12 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] (duration: 10m 59s)
* 22:06 musikanimal@deploy1003: kemayo, musikanimal: Rolling back deployment
* 22:05 musikanimal@deploy1003: kemayo, musikanimal: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:01 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]]
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:00 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:52 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:47 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:47 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] (duration: 13m 23s)
* 21:46 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:46 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2006.codfw.wmnet
* 21:42 derenrich@deploy1003: derenrich: Continuing with deployment
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1348
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1348
* 21:39 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1348
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:39 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:37 derenrich@deploy1003: derenrich: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:35 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:34 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1348
* 21:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 21:33 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1348.eqiad.wmnet
* 21:33 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]]
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1348.eqiad.wmnet
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1348.eqiad.wmnet
* 21:31 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2006.codfw.wmnet
* 21:31 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] (duration: 09m 45s)
* 21:27 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:26 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]]
* 21:19 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:19 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2006.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 21:17 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] (duration: 13m 54s)
* 21:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2005.codfw.wmnet with OS bookworm
* 21:12 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:07 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdeb
* 21:03 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]]
* 20:54 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:52 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] (duration: 23m 49s)
* 20:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:47 jdrewniak@deploy1003: jdrewniak, milazg: Continuing with deployment
* 20:34 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1346.eqiad.wmnet
* 20:32 jdrewniak@deploy1003: jdrewniak, milazg: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2005.codfw.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:28 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]]
* 20:27 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2005.codfw.wmnet
* 20:26 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:24 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:22 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] (duration: 11m 24s)
* 20:17 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 20:15 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2005.codfw.wmnet
* 20:14 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:12 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 20:11 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:07 jdrewniak@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.18,1.47.0-wmf.19,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted
* 20:05 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:02 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:01 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2005.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 19:53 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:50 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1346
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 19:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1346.eqiad.wmnet
* 19:21 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 thcipriani: Gerrit downtime incoming for upgrade
* 19:17 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hadoop.roll-restart-workers (exit_code=0) restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 dzahn@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:30:00 on gerrit.wikimedia.org with reason: maintenance upgrade
* 19:16 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on gerrit2003.wikimedia.org with reason: maintenance upgrade
* 19:04 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 18:21 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e] (duration: 00m 59s)
* 18:20 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e]
* 18:19 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e] (duration: 05m 13s)
* 18:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 18:14 otto@deploy1003: Started deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e]
* 18:13 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e] (duration: 00m 39s)
* 18:13 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e]
* 18:13 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 18:12 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 18:11 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 18:11 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit2002.wikimedia.org with reason: maintenance upgrade
* 18:11 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:11 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 18:10 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit1003.wikimedia.org with reason: maintenance upgrade
* 18:09 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 18:08 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 18:06 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 16:40 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 16:35 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 16:33 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]] synced to the te
* 16:28 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 15:33 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4
* 15:04 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] (duration: 08m 08s)
* 15:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host clouddumps1001.wikimedia.org with OS bookworm
* 14:59 samtar@deploy1003: samtar: Continuing with deployment
* 14:58 samtar@deploy1003: samtar: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:56 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]]
* 14:40 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:07 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 13:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host clouddumps1001.wikimedia.org with OS bookworm
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:55 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:51 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:42 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:41 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:48 klausman@dns1004: END - running authdns-update
* 12:46 klausman@dns1004: START - running authdns-update
* 12:35 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 12:35 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 12:05 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning x4
* 12:01 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2005.codfw.wmnet
* 11:55 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2005.codfw.wmnet
* 11:54 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1144.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:52 cgoubert@dns1004: END - running authdns-update
* 11:49 cgoubert@dns1004: START - running authdns-update
* 11:31 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2004.codfw.wmnet
* 11:25 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2004.codfw.wmnet
* 11:24 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:16 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1200.eqiad.wmnet
* 11:16 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1200.eqiad.wmnet
* 11:04 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1200.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:04 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1199.eqiad.wmnet
* 11:03 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1199.eqiad.wmnet
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:42 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1199.eqiad.wmnet with reason: Upgrading RAID firmware
* 10:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1024.eqiad.wmnet with reason: Cloning x4
* 10:00 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 09:56 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 09:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1024.eqiad.wmnet
* 09:45 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 09:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 09:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:04 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:51 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:46 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning x4
* 08:43 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:34 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] (duration: 09m 56s)
* 08:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:30 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:29 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:28 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:24 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]]
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 08:09 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 07:58 XioNoX: netflow1004:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:56 XioNoX: netflow2005:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 07:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 07:03 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:59 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:43 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:42 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:39 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 06:38 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 06:37 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 05:04 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 05:03 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tool-server' for release 'main' .
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 38s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1345.eqiad.wmnet
* 00:11 dzahn@dns1004: END - running authdns-update
* 00:08 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 00:08 dzahn@dns1004: START - running authdns-update
== 2026-09-09 ==
* 23:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:29 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1345
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 23:28 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1345.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1344.eqiad.wmnet
* 23:22 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] (duration: 11m 15s)
* 23:18 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 23:16 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 23:11 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]]
* 22:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:51 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2004.codfw.wmnet with OS bookworm
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1344
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1344
* 22:37 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1344
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1344
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 22:32 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] (duration: 10m 21s)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1344.eqiad.wmnet
* 22:27 derenrich@deploy1003: derenrich, egardner: Continuing with deployment
* 22:26 derenrich@deploy1003: derenrich, egardner: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:24 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:22 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]]
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:19 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] (duration: 13m 40s)
* 22:16 derenrich@deploy1003: derenrich: Rolling back deployment
* 22:10 derenrich@deploy1003: derenrich: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:05 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]]
* 22:03 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:44 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:40 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:36 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1343
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1343
* 21:27 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1343
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:27 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:23 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:22 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1343
* 21:22 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] (duration: 12m 29s)
* 21:21 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 21:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1343.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1342.eqiad.wmnet
* 21:17 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:14 jforrester@deploy1003: jforrester: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:09 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]]
* 21:08 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:07 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 21:07 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:06 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2004.codfw.wmnet
* 20:59 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:57 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:56 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2004.codfw.wmnet
* 20:53 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:53 bking@cumin2003: END (ERROR) - Cookbook sre.presto.roll-restart-workers (exit_code=97) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:53 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:50 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* {{safesubst:SAL entry|1=20:45 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2}}
* 20:42 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* 20:41 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 20:40 sbassett@deploy1003: aranyap, sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:39 sbassett@deploy1003: aranyap, sbassett: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "Filter}}
* 20:35 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* {{safesubst:SAL entry|1=20:34 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "}}
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:28 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:24 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 20:24 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1342
* 20:23 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 20:23 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1342.eqiad.wmnet
* 20:23 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1342.eqiad.wmnet
* 20:22 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1342.eqiad.wmnet
* 20:15 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1027.eqiad.wmnet with OS bookworm
* 19:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:41 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1027.eqiad.wmnet with OS bookworm
* 19:36 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:28 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:26 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:19 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 19:19 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:06 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 19:05 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 19:03 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 18:59 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 18:23 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:06 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1026.eqiad.wmnet with OS bookworm
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 18:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 17:54 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 17:49 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 17:45 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1026.eqiad.wmnet with OS bookworm
* 17:31 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:29 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 17:27 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 17:12 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 17:04 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 16:46 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] (duration: 09m 28s)
* 16:41 urbanecm@deploy1003: migr, urbanecm: Continuing with deployment
* 16:41 urbanecm@deploy1003: migr, urbanecm: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:36 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]]
* 16:28 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:25 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 15:55 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 15:54 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 15:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] (duration: 09m 43s)
* 15:41 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 15:40 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:36 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]]
* 15:17 urbanecm: Delete all running periodic jobs starting with `growthexperiments-refreshlinkrecommendations-*` (to pick up new configuration; [[phab:T392944|T392944]])
* 15:08 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:06 moritzm: installing grub2 bugfix updates from Bookworm point release
* 15:04 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp6008.drmrs.wmnet
* 15:01 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:42 hnowlan: half concurrency for parsoidCachePrewarm RecordLintJob and refreshLinks in jobqueue, eqiad & codfw
* 14:35 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 14:34 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:32 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 14:20 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:07 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2013.codfw.wmnet with OS trixie
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 14:06 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 13:55 btullis@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'.
* 13:53 btullis@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'.
* 13:43 btullis@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 13:42 btullis@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 13:29 moritzm: pruned obsolete Bullseye image dispatch from the docker registry [[phab:T416452|T416452]]
* 13:28 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:26 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b7-eqiad
* 13:25 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 13:24 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:22 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 13:22 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:17 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a4-eqiad
* 13:17 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] (duration: 10m 15s)
* 13:14 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 13:11 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 13:10 sbisson@deploy1003: sbisson: Continuing with deployment
* 13:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:04 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]]
* 13:02 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on ldap-rw[1001,2001].wikimedia.org with reason: work in progress
* 12:49 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 12:48 btullis@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
* 12:46 btullis@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
* 12:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] (duration: 14m 39s)
* 12:41 ladsgroup@deploy1003: tryvix1509, ladsgroup: Continuing with deployment
* 12:35 ladsgroup@deploy1003: tryvix1509, ladsgroup: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:31 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]]
* 12:16 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 12:16 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 11:53 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] (duration: 21m 58s)
* 11:48 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 11:35 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:31 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]]
* 10:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Repooling db2196
* 10:47 moritzm: pruned obsolete Bullseye images nodejs12-slim/nodejs12-devel/nodejs14-slim/nodejs16-slim from the docker registry [[phab:T416452|T416452]]
* 10:43 moritzm: installing Bird security updates
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1260: Repooling after cloning
* 10:09 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Repooling db2196
* 10:07 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 10:06 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 09:55 moritzm: pruned obsolete Bullseye images openjdk-8-jdk/openjdk-8-jre/openjdk-11-jre/openjdk-11-jdk from the docker registry [[phab:T416452|T416452]]
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1260: Repooling after cloning
* 09:52 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 09:52 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 09:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:28 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:27 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:03 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 09:02 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 09:01 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 16 hosts with reason: upgrade ssw1-a1-eqiad
* 08:58 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 22 hosts with reason: upgrade ssw1-a1-eqiad
* 08:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 08:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 08:49 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 08:49 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 08:48 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 08:48 cmooney@cumin1004: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 08:36 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning sanitarium
* 08:30 brouberol@dns1004: END - running authdns-update
* 08:28 moritzm: pruned obsolete Bullseye image golang1.15 from the docker registry [[phab:T416452|T416452]]
* 08:28 brouberol@dns1004: START - running authdns-update
* 08:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1260: Needs to clone another host from this one
* 08:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1260: Needs to clone another host from this one
* 08:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1260.eqiad.wmnet with reason: Cloning sanitarium
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:00 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 07:40 chlod: UTC morning backport window done
* 07:37 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] (duration: 21m 36s)
* 07:32 chlod@deploy1003: chlod, hamishz: Continuing with deployment
* 07:20 chlod@deploy1003: chlod, hamishz: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 45s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1025.eqiad.wmnet with OS bookworm
== 2026-09-08 ==
* 23:51 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:48 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:39 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1313.eqiad.wmnet
* 23:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:25 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 23:19 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:19 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:15 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:05 Amir1: dropped 57 tables on db1260 ([[phab:T437278|T437278]])
* 23:03 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 23:03 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:02 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 22:57 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1313
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1313
* 22:37 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1313
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:37 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:33 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 22:33 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1313
* 22:32 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 22:32 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1313.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1306.eqiad.wmnet
* 22:27 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6008.drmrs.wmnet with OS trixie
* 22:18 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 22:03 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 22:01 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] (duration: 09m 53s)
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 21:58 Amir1: drop links tables from db2210 ([[phab:T437278|T437278]])
* 21:57 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 21:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:56 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:52 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 21:51 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]]
* 21:51 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:48 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:45 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] (duration: 05m 27s)
* 21:43 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Continuing with deployment
* 21:40 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:39 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]]
* 21:35 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1024.eqiad.wmnet with OS bookworm
* 21:33 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp6008.drmrs.wmnet with OS trixie
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1306
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1306
* 21:30 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1306
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:24 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] (duration: 09m 12s)
* 21:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 21:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1306
* 21:19 reedy@deploy1003: reedy: Continuing with deployment
* 21:19 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:19 reedy@deploy1003: reedy: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1306.eqiad.wmnet
* 21:15 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:15 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]]
* {{safesubst:SAL entry|1=21:10 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out}}
* 21:05 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1024.eqiad.wmnet with OS bookworm
* 21:05 sbassett@deploy1003: sbassett: Continuing with deployment
* 21:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 21:04 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=21:03 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out non-http(s) lice}}
* 20:59 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=20:58 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out n}}
* 20:53 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 20:50 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1305.eqiad.wmnet
* 20:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 20:34 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 20:28 sbassett@deploy1003: sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:27 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-http(s) license}}
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1023.eqiad.wmnet with OS bookworm
* {{safesubst:SAL entry|1=20:23 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-}}
* 20:15 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] (duration: 10m 16s)
* 20:13 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:10 aaron@deploy1003: aaron: Continuing with deployment
* 20:09 aaron@deploy1003: aaron: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 20:05 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:05 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]]
* 20:01 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1023.eqiad.wmnet with OS bookworm
* 19:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1305
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1305
* 19:43 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1305
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:43 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:39 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 19:39 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1305
* 19:38 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 19:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1305.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1275.eqiad.wmnet
* 19:23 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 19:17 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 18:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:53 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1275
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1275
* 18:33 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1275
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1275
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 18:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1275.eqiad.wmnet
* 18:18 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 17:43 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy1003.eqiad.wmnet
* 17:36 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet
* 17:36 cdobbins@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-ntp (exit_code=0) rolling restart_daemons on A:dnsbox
* 17:30 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy1003.eqiad.wmnet
* 17:25 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet
* 17:15 swfrench@deploy1003: Finished scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup (duration: 04m 18s)
* 17:11 Amir1: dropping links tables from db1247 (s4 replica) - ([[phab:T437278|T437278]])
* 17:10 swfrench@deploy1003: Started scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup
* 16:51 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] (duration: 10m 19s)
* 16:46 zabe@deploy1003: zabe: Continuing with deployment
* 16:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]]
* 16:29 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 jhancock@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:12 jhancock@cumin2003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 16:08 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2003.codfw.wmnet
* 15:58 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:55 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2003.codfw.wmnet
* 15:54 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:44 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2002.codfw.wmnet
* 15:19 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2002.codfw.wmnet
* 14:55 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:46 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2001.codfw.wmnet
* 14:44 topranks: shutdown et-1/1/5 on cr1-codfw to shift traffic off ssw1-a1-codfw
* 14:43 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: upgrade ssw1-a1-eqiad
* 14:34 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2001.codfw.wmnet
* 14:33 btullis@cumin1003: END (ERROR) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=97) rolling reboot on A:cephosd-codfw
* 14:30 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet
* 14:28 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet
* 14:22 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet
* 14:13 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --olderThan {{Gerrit|1788220800}} --verbose # [[phab:T437158|T437158]]
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet
* 14:12 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet
* 14:03 topranks: drain traffic from ssw1-a1-codfw before JunOS upgrade [[phab:T426197|T426197]]
* 14:02 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet
* 13:58 cgoubert@deploy1003: helmfile [staging-codfw] DONE helmfile.d/services/mw-debug: apply
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet
* 13:57 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet
* 13:56 cgoubert@deploy1003: helmfile [staging-codfw] START helmfile.d/services/mw-debug: apply
* 13:55 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:54 stran@deploy1003: mwscript-k8s job started: foreachwikiindblist checkuser-suggested-investigations extensions/CheckUser/maintenance/populateSiCaseProperties.php # [[phab:T435066|T435066]]
* 13:52 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:51 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet
* 13:50 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:46 cdobbins@cumin1003: START - Cookbook sre.dns.roll-restart-ntp rolling restart_daemons on A:dnsbox
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet
* 13:46 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet
* 13:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2012.codfw.wmnet with OS bookworm
* 13:44 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] (duration: 34m 00s)
* 13:40 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet
* 13:37 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet
* 13:35 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet
* 13:32 stran@deploy1003: stran: Continuing with deployment
* 13:29 stran@deploy1003: stran: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet
* 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:21 moritzm: installing qemu security updates
* 13:18 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet
* 13:16 ayounsi@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'email' for AS: 139628
* 13:15 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 139628
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet
* 13:13 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet
* 13:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 2519
* 13:11 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 2519
* 13:10 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 14593
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]]
* 13:09 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:08 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 14593
* 13:06 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet
* 13:06 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2012.codfw.wmnet with OS bookworm
* 13:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2012.codfw.wmnet
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2012.codfw.wmnet
* 13:05 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:04 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet
* 13:01 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet
* 12:58 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 34655
* 12:58 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 34655
* 12:56 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2012.codfw.wmnet
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'clear' for AS: 35320
* 12:55 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'clear' for AS: 35320
* 12:55 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-codfw
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-codfw
* 12:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2011.codfw.wmnet with OS bookworm
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-by27-esams
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-by27-esams
* 12:54 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-esams
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-esams
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-esams
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-esams
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-eqsin
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-eqsin
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-ulsfo
* 12:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f3-eqiad
* 12:50 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f3-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-c8-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-c8-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:44 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:43 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet
* 12:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-eqiad
* 12:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f1-eqiad
* 12:41 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f1-eqiad
* 12:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-d5-eqiad
* 12:40 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-d5-eqiad
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet
* 12:38 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet
* 12:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:32 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet
* 12:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad
* 12:11 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2012.codfw.wmnet
* 12:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2011.codfw.wmnet with OS bookworm
* 12:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2011.codfw.wmnet
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2011.codfw.wmnet
* 12:09 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:06 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] (duration: 09m 54s)
* 12:01 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2011.codfw.wmnet
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Continuing with deployment
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2010.codfw.wmnet with OS bookworm
* 11:56 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]]
* 11:46 marostegui@dns1004: END - running authdns-update
* 11:44 marostegui@dns1004: START - running authdns-update
* 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:40 Amir1: dropping unneeded tables from x4 - db1260 ([[phab:T437278|T437278]])
* 11:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:23 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2011.codfw.wmnet
* 11:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2010.codfw.wmnet with OS bookworm
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2010.codfw.wmnet
* 11:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2010.codfw.wmnet
* 11:20 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2010.codfw.wmnet
* 11:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2009.codfw.wmnet with OS bookworm
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:43 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] (duration: 10m 57s)
* 10:39 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:38 samtar@deploy1003: samtar: Continuing with deployment
* 10:37 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:37 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:36 samtar@deploy1003: samtar: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:34 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:32 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]]
* 10:30 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2009.codfw.wmnet with OS bookworm
* 10:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2009.codfw.wmnet
* 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:18 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2009.codfw.wmnet
* 10:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2008.codfw.wmnet with OS bookworm
* 10:07 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:05 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:01 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:45 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] (duration: 13m 15s)
* 09:45 ayounsi@dns1004: END - running authdns-update
* 09:44 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:43 ayounsi@dns1004: START - running authdns-update
* 09:39 zabe@deploy1003: zabe: Continuing with deployment
* 09:37 zabe@deploy1003: zabe: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2008.codfw.wmnet with OS bookworm
* 09:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:32 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]]
* 09:32 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2008.codfw.wmnet
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2008.codfw.wmnet
* 09:31 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:29 XioNoX: remove GRE tunnels eqiad-drmrs eqdfw-ulsfo
* 09:23 moritzm: installing rsync security updates
* 09:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2007.codfw.wmnet with OS bookworm
* 09:22 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2008.codfw.wmnet
* 09:11 marostegui@cumin1003: dbctl commit (dc=all): 'Make x4 and s4 RW again [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96393 and previous config saved to /var/cache/conftool/dbconfig/20260908-091121-marostegui.json
* 09:07 marostegui@cumin1003: dbctl commit (dc=all): 'Remove old s4 masters from x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96392 and previous config saved to /var/cache/conftool/dbconfig/20260908-090749-marostegui.json
* 09:05 marostegui@cumin1003: dbctl commit (dc=all): 'Set x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96391 and previous config saved to /var/cache/conftool/dbconfig/20260908-090517-marostegui.json
* 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Set s4 commons to read-only for maintenance [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96389 and previous config saved to /var/cache/conftool/dbconfig/20260908-090228-marostegui.json
* 09:02 marostegui: Starting x4 split from s4, RO time on commons needed [[phab:T404715|T404715]]
* 09:00 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 08:58 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2008.codfw.wmnet
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:43 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 32 hosts with reason: x4 split
* 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2007.codfw.wmnet with OS bookworm
* 08:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:37 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:37 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2007.codfw.wmnet
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2007.codfw.wmnet
* 08:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:34 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:29 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:26 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2007.codfw.wmnet
* 08:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2006.codfw.wmnet with OS bookworm
* 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:50 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2007.codfw.wmnet
* 07:43 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2006.codfw.wmnet with OS bookworm
* 07:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2006.codfw.wmnet
* 07:37 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:30 denisse: Add grafana-plugins 0.15 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 07:29 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:27 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 07:27 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:27 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:22 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 07:22 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 07:18 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:14 jmm@dns1004: END - running authdns-update
* 07:13 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96388 and previous config saved to /var/cache/conftool/dbconfig/20260908-071308-marostegui.json
* 07:12 jmm@dns1004: START - running authdns-update
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96387 and previous config saved to /var/cache/conftool/dbconfig/20260908-071216-marostegui.json
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96386 and previous config saved to /var/cache/conftool/dbconfig/20260908-071159-marostegui.json
* 05:07 denisse: Add grafana-plugins 0.10 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.16 (duration: 02m 27s)
* 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]] (duration: 36m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 41s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-07 ==
* 21:52 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] (duration: 11m 00s)
* 21:47 zabe@deploy1003: zabe: Continuing with deployment
* 21:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]]
* 21:37 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] (duration: 09m 34s)
* 21:33 zabe@deploy1003: zabe: Continuing with deployment
* 21:32 zabe@deploy1003: zabe: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:28 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]]
* 21:03 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] (duration: 10m 27s)
* 20:58 zabe@deploy1003: zabe: Continuing with deployment
* 20:57 zabe@deploy1003: zabe: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:52 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]]
* 20:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set weight of db1261 to zero in s4 ([[phab:T437108|T437108]])', diff saved to https://phabricator.wikimedia.org/P96385 and previous config saved to /var/cache/conftool/dbconfig/20260907-203804-ladsgroup.json
* 20:23 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] (duration: 09m 28s)
* 20:19 zabe@deploy1003: zabe: Continuing with deployment
* 20:18 zabe@deploy1003: zabe: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:14 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]]
* 20:12 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] (duration: 10m 06s)
* 20:07 zabe@deploy1003: zabe, daimona: Continuing with deployment
* 20:06 zabe@deploy1003: zabe, daimona: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:02 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]]
* 19:59 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] (duration: 11m 27s)
* 19:55 zabe@deploy1003: zabe: Continuing with deployment
* 19:52 zabe@deploy1003: zabe: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:48 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]]
* 19:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] (duration: 11m 40s)
* 19:26 zabe@deploy1003: zabe: Continuing with deployment
* 19:23 zabe@deploy1003: zabe: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:19 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]]
* 19:07 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] (duration: 14m 17s)
* 19:00 zabe@deploy1003: zabe: Continuing with deployment
* 18:57 zabe@deploy1003: zabe: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:53 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]]
* 18:33 zabe@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 18:32 zabe@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 18:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] (duration: 09m 12s)
* 18:27 zabe@deploy1003: zabe: Continuing with deployment
* 18:26 zabe@deploy1003: zabe: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:22 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]]
* 16:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2005.codfw.wmnet with OS bookworm
* 16:01 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] (duration: 10m 22s)
* 15:59 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 15:57 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 15:56 zabe@deploy1003: zabe: Continuing with deployment
* 15:55 zabe@deploy1003: zabe: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:51 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]]
* 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:47 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 15:46 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 15:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:44 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 15:44 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 15:27 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2005.codfw.wmnet with OS bookworm
* 15:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2005.codfw.wmnet
* 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:11 moritzm: installing rsync security updates
* 15:04 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1341.eqiad.wmnet
* 15:03 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2004.codfw.wmnet with OS bookworm
* 15:00 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:58 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 14:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:40 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 14:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 14:35 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:32 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:29 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:25 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 14:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1228: Repooling db1228 into s4
* 14:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2004.codfw.wmnet with OS bookworm
* 14:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Repooling after cloning
* 14:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:19 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2004.codfw.wmnet
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2004.codfw.wmnet
* 14:18 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:14 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1074.eqiad.wmnet
* 14:14 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:13 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:13 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* {{safesubst:SAL entry|1=14:11 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mul}}
* 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2004.codfw.wmnet
* 14:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1074.eqiad.wmnet
* 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1073.eqiad.wmnet
* 14:07 krinkle@deploy1003: krinkle: Continuing with deployment
* {{safesubst:SAL entry|1=14:04 krinkle@deploy1003: krinkle: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with multiple properties}}
* 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1073.eqiad.wmnet
* 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1072.eqiad.wmnet
* 14:01 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1341
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1341
* 14:01 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 14:00 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1341
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* 13:59 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* {{safesubst:SAL entry|1=13:59 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mult}}
* 13:59 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2004.codfw.wmnet
* 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2003.codfw.wmnet with OS bookworm
* 13:55 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 13:55 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1072.eqiad.wmnet
* 13:55 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudvirt1067.eqiad.wmnet
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1341
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 13:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1341.eqiad.wmnet
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2008.wikimedia.org
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2008.wikimedia.org with OS trixie
* 13:39 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1067.eqiad.wmnet
* 13:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1066.eqiad.wmnet
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:37 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1228: Repooling db1228 into s4
* 13:36 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Repooling after cloning
* 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1066.eqiad.wmnet
* 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1065.eqiad.wmnet
* 13:28 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:27 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1065.eqiad.wmnet
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:25 moritzm: installing openssh security updates
* 13:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:24 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1340.eqiad.wmnet
* 13:23 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] (duration: 10m 06s)
* 13:17 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2003.codfw.wmnet with OS bookworm
* 13:17 stran@deploy1003: stran: Continuing with deployment
* 13:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:16 stran@deploy1003: stran: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:16 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:15 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 13:12 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]]
* 13:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2003.codfw.wmnet
* 13:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2003.codfw.wmnet
* 13:03 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2008.wikimedia.org with OS trixie
* 13:03 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:02 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:02 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:01 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:00 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2003.codfw.wmnet
* 12:58 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:54 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 12:54 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2008.wikimedia.org
* 12:47 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2003.codfw.wmnet
* 12:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2002.codfw.wmnet with OS bookworm
* 12:29 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 12:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2007.wikimedia.org
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2007.wikimedia.org with OS trixie
* 12:14 elukey: moved most of the Docker Registry's prefixes to a new internal S3 backend. For any docker pull failure that worked in the past, please ping me or drop a note in [[phab:T435499|T435499]] or contact the oncall SREs
* 12:07 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2002.codfw.wmnet with OS bookworm
* 12:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:02 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2002.codfw.wmnet
* 11:54 jmm@dns1004: END - running authdns-update
* 11:52 jmm@dns1004: START - running authdns-update
* 11:46 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2002.codfw.wmnet
* 11:46 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2007.wikimedia.org with OS trixie
* 11:46 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:46 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:41 moritzm: installing bash updates from bookworm point release
* 11:39 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:39 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2007.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts ldap-replica1006.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:35 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:32 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2001.codfw.wmnet with OS bookworm
* 11:28 jmm@cumin1004: START - Cookbook sre.hosts.decommission for hosts ldap-replica1006.wikimedia.org
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 11:18 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] (duration: 14m 08s)
* 11:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 11:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet
* 11:11 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:11 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:10 zabe@deploy1003: zabe: Continuing with deployment
* 11:10 zabe@deploy1003: zabe: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 11:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet
* 11:07 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet
* 11:07 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet
* 11:06 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:06 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:04 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]]
* 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 10:58 btullis@deploy1003: Finished scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]] (duration: 35m 20s)
* 10:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:53 jmm@dns1004: END - running authdns-update
* 10:51 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:50 jmm@dns1004: START - running authdns-update
* 10:47 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=frwiki # [[phab:T436659|T436659]]
* 10:40 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=hrwiki # [[phab:T436659|T436659]]
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:39 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 10:37 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 10:33 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 10:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2001.codfw.wmnet with OS bookworm
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1340
* 10:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1340.eqiad.wmnet
* 10:26 btullis@deploy1003: Started scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]]
* 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2001.codfw.wmnet
* 10:18 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2001.codfw.wmnet
* 10:13 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw
* 10:12 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:10 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:09 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:08 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:07 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:07 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:02 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:00 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:59 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:57 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/services/miscweb: apply
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/services/miscweb: apply
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:52 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:52 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] (duration: 10m 11s)
* 09:51 aokoth@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 09:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 09:49 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 09:49 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 09:48 aokoth@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 09:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:45 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:44 aokoth@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 09:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]]
* 09:41 aokoth@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 09:38 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Repooling after switchover
* 09:23 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw
* 09:15 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] (duration: 20m 12s)
* 09:12 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 139009
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1005.wikimedia.org
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1005.wikimedia.org with OS trixie
* 09:10 moritzm: rebuild software RAID following disk replacement [[phab:T437036|T437036]]
* 09:10 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 139009
* 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1022.eqiad.wmnet with OS bookworm
* 09:08 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 09:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:03 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet
* 09:02 moritzm: installing giflib security updates
* 09:01 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:59 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 08:57 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 08:56 jmm@cumin1004: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet
* 08:56 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:54 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]]
* 08:52 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:52 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:49 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 08:48 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:42 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:41 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:40 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:40 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96355 and previous config saved to /var/cache/conftool/dbconfig/20260907-083904-marostegui.json
* 08:38 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1005.wikimedia.org with OS trixie
* 08:38 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2157 to s5 primary [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96354 and previous config saved to /var/cache/conftool/dbconfig/20260907-083825-marostegui.json
* 08:38 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 marostegui: Starting s5 codfw failover from db2213 to db2157 - [[phab:T437188|T437188]]
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2157 with weight 0 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96353 and previous config saved to /var/cache/conftool/dbconfig/20260907-083448-marostegui.json
* 08:34 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s5 [[phab:T437188|T437188]]
* 08:28 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:28 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1005.wikimedia.org
* 08:22 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:20 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1022.eqiad.wmnet with OS bookworm
* 08:03 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:02 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 08:02 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:00 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 07:57 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 07:56 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1241.eqiad.wmnet with reason: Cloning
* 07:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1241: Cloning
* 07:55 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1241: Cloning
* 07:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Cloning
* 07:54 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Cloning
* 07:51 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 07:50 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 07:47 kartik@deploy1003: Finished scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] (duration: 41m 51s)
* 07:46 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 07:45 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 07:34 kartik@deploy1003: abi, kartik: Continuing with deployment
* 07:23 kartik@deploy1003: abi, kartik: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:05 kartik@deploy1003: Started scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]]
* 06:14 moritzm: installing Chromium security updates
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 08m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-06 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 25s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-05 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 26s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-04 ==
* 22:07 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 21:48 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:42 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:30 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:42 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:40 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 19:27 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:19 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:13 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:12 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host sretest2013
* 19:11 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host sretest2013
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:11 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:07 jhancock@cumin1003: START - Cookbook sre.dns.netbox
* 18:18 inflatador: bking@clouddumps100[12] `systemctl reset-failed` to quash alerts until https://w.wiki/UBje . The systemd timer should try again tomorrow
* 17:27 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b8-eqiad
* 17:27 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b8-eqiad
* 16:37 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:33 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-b7-eqiad
* 16:33 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 16:05 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b6-eqiad
* 16:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b6-eqiad
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1339.eqiad.wmnet
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 15:50 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:49 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:47 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:43 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 15:10 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b5-eqiad
* 15:09 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b5-eqiad
* 14:46 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1045.eqiad.wmnet
* 14:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1003.eqiad.wmnet with OS trixie
* 14:38 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 14:37 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b4-eqiad
* 14:37 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b4-eqiad
* 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339
* 14:32 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339
* 14:31 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 14:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet
* 14:26 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:24 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:21 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b3-eqiad
* 14:21 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b3-eqiad
* 14:17 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:17 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:04 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1003.eqiad.wmnet with OS trixie
* 13:54 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b2-eqiad
* 13:53 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b2-eqiad
* 13:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1002.eqiad.wmnet with OS trixie
* 13:18 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a4-eqiad
* 13:12 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:12 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b1-eqiad
* 13:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b1-eqiad
* 12:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1002.eqiad.wmnet with OS trixie
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow3004.esams.wmnet with OS trixie
* 12:33 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:28 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 12:15 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 12:11 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a7-eqiad
* 12:11 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a7-eqiad
* 12:01 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow3004.esams.wmnet with OS trixie
* 11:47 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a6-eqiad
* 11:47 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a6-eqiad
* 11:36 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki2003.codfw.wmnet
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki2003.codfw.wmnet with OS trixie
* 11:19 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:13 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a5-eqiad
* 11:06 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a5-eqiad
* 10:52 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki2003.codfw.wmnet with OS trixie
* 10:50 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:50 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:44 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 10:44 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki2003.codfw.wmnet
* 10:29 btullis@deploy1003: Finished scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli (duration: 41m 14s)
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 10:00 marostegui@cumin1003: Removing db1182 from zarcillo [[phab:T434869|T434869]]
* 10:00 marostegui@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db1182.eqiad.wmnet
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:57 btullis@deploy1003: Started scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli
* 09:53 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:53 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.decommission (exit_code=1)
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1182.eqiad.wmnet
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:50 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:46 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:45 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1182 from dbctl [[phab:T434869|T434869]]', diff saved to https://phabricator.wikimedia.org/P96346 and previous config saved to /var/cache/conftool/dbconfig/20260904-094527-marostegui.json
* 09:41 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki1003.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki1003.eqiad.wmnet with OS trixie
* 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1182: Decommissioning
* 09:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1182: Decommissioning
* 09:23 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 09:17 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow5003.eqsin.wmnet with OS trixie
* 09:02 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki1003.eqiad.wmnet with OS trixie
* 09:00 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 09:00 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:55 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:55 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki1003.eqiad.wmnet
* 08:54 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:48 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:45 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:40 btullis@deploy1003: Finished scap sync-world: Trying again for [[phab:T436913|T436913]] (duration: 34m 26s)
* 08:35 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2003.codfw.wmnet
* 08:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw2001.wikimedia.org with OS trixie
* 08:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4 days, 0:00:00 on db2196.codfw.wmnet with reason: Host crashed
* 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2003.codfw.wmnet
* 08:24 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:20 elukey@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:12 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:08 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:07 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 08:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:04 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 07:57 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow5003.eqsin.wmnet with OS trixie
* 07:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:55 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw2001.wikimedia.org with OS trixie
* 07:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:51 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:50 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie
* 07:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2196: down
* 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2196: down
* 06:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:52 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:40 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:28 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:21 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 06:20 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 39s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 01:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:02 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:59 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:46 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 00:44 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:39 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:23 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
== 2026-09-03 ==
* 21:46 tsev@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T435363|T435363]]
* 21:03 eevans@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:55 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] (duration: 12m 24s)
* 20:52 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:52 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1021.eqiad.wmnet with OS bookworm
* 20:50 arlolra@deploy1003: arlolra, tsev: Continuing with deployment
* 20:46 arlolra@deploy1003: arlolra, tsev: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:44 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1047.eqiad.wmnet
* 20:42 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]]
* 20:41 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a3-eqiad
* 20:40 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a3-eqiad
* 20:39 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] (duration: 10m 23s)
* 20:34 arlolra@deploy1003: arlolra, jgiannelos: Continuing with deployment
* 20:33 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1047.eqiad.wmnet
* 20:33 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:32 arlolra@deploy1003: arlolra, jgiannelos: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:31 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1046.eqiad.wmnet
* 20:28 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]]
* 20:23 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] (duration: 13m 41s)
* 20:20 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1046.eqiad.wmnet
* 20:16 catrope@deploy1003: catrope: Continuing with deployment
* 20:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1021.eqiad.wmnet with OS bookworm
* 20:13 catrope@deploy1003: catrope: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1021.eqiad.wmnet
* 20:09 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]]
* 20:00 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1021.eqiad.wmnet
* 19:49 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 19:19 swfrench@deploy1003: Finished scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]] (duration: 02m 59s)
* 19:16 swfrench@deploy1003: Started scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]]
* 19:01 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:01 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 18:57 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:57 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:34 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:20 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 17:55 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:55 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:53 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:52 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:46 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 17:46 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 17:45 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 17:44 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 17:44 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:40 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:39 ryankemper: [WDQS] Service looks healthy again, CPU load and thread count have dropped considerably over the last hour
* 17:39 ryankemper: [[phab:T421642|T421642]] [WDQS] requestctl changes: `2026-09-03 16:23-17:33` UTC: added hard-deny pair `cache-text/wdqs_futile_sparql_sep_2026_deny(+_bots)`; extended pattern `ua/wdqs_heavy_sparql_bots_2026` and added default-scope twin `wdqs_heavy_sparql_bots_jul_2026_ratelimit_default`; added ipblock `abuse/wdqs_sparql_scanners_sep_2026` + throttle `wdqs_sparql_scanners_sep_2026_ratelimit` (needed manual `requestctl update-provenance-map`)
* 17:37 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:37 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:31 andrew@cumin2003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudcephosd1045.eqiad.wmnet
* 17:30 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:30 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:28 dancy@deploy1003: Installation of scap version "4.289.0" completed for 3 hosts
* 17:26 dancy@deploy1003: Installing scap version "4.289.0" for 3 host(s)
* 17:24 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a2-eqiad
* 17:24 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:22 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:22 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a2-eqiad
* 17:14 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 17:10 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:10 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:09 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 17:01 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:59 andrew@cumin2003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 16:58 dancy: Running scap clean-images on deploy1003
* 16:52 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:50 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:39 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:14 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs2021.codfw.wmnet,service=wdqs-main
* 16:14 ryankemper: [[phab:T430880|T430880]] Stumbled across `wdqs2021` listed as inactive, looks like it was never fully re-pooled after a data xfer. Pooled.
* 16:12 btullis@deploy1003: Finished deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2] (duration: 00m 38s)
* 16:12 btullis@deploy1003: Started deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2]
* 16:12 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:07 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs101[1-4].eqiad.wmnet
* 16:03 btullis@deploy1003: Started scap sync-world: Rebuilding to pick up new version of dump scripts in mediawiki-cli for [[phab:T436913|T436913]]
* 16:01 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]] (duration: 09m 29s)
* 15:51 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=urldownloader[12]00[56].wikimedia.org [reason: depooling urldownloader trixie nodes]
* 15:51 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]]
* 15:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:29 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:24 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:24 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:18 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:15 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:09 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:05 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 15:04 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:02 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 15:00 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader
* 14:58 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw
* 14:53 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-ulsfo
* 14:49 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 14:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2002.codfw.wmnet
* 14:45 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:39 sukhe: sudo cumin "A:cp-text" "run-puppet-agent --enable 'merging CR 1334855'": [[phab:T425441|T425441]]
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2002.codfw.wmnet
* 14:32 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:30 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:29 sukhe: sudo cumin "A:cp-text" "disable-puppet 'merging CR 1334855'" [[phab:T425441|T425441]]
* 14:27 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-ulsfo
* 14:22 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:21 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:19 arnaudb@dns1006: END - running authdns-update
* 14:18 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1074
* 14:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1074
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-codfw
* 14:17 arnaudb@dns1006: START - running authdns-update
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:16 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:16 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1073
* 14:15 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1073
* 14:12 ayounsi@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host netflow2004.codfw.wmnet with OS trixie
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:10 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:08 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:08 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:07 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] (duration: 09m 36s)
* 14:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:05 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:03 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 14:03 samtar@deploy1003: btullis, samtar: Continuing with deployment
* 14:02 samtar@deploy1003: btullis, samtar: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:01 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:58 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]]
* 13:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:56 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:55 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:52 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqiad and A:durum
* 13:52 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-codfw
* 13:51 moritzm: installing sqlite3 security updates
* 13:51 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:51 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqiad and A:durum
* 13:49 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-codfw and A:durum
* 13:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:47 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-codfw and A:durum
* 13:47 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-esams
* 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:44 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:43 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:43 ayounsi@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:42 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] (duration: 13m 50s)
* 13:41 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:40 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1072
* 13:40 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:39 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-esams and A:durum
* 13:38 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1072
* 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:38 samtar@deploy1003: samtar, thiemowmde: Continuing with deployment
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-esams and A:durum
* 13:37 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:33 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:33 samtar@deploy1003: samtar, thiemowmde: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:32 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqsin and A:durum
* 13:31 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqsin and A:durum
* 13:28 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]]
* 13:28 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 13:27 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:24 ayounsi@cumin1004: START - Cookbook sre.hosts.reimage for host netflow2004.codfw.wmnet with OS trixie
* 13:24 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:22 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 13:22 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-esams
* 13:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:15 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:15 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 moritzm: installing bash updates from trixie point release
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 13:14 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1067
* 13:13 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1067
* 13:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2901: Test
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:09 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:09 moritzm: installing libxslt bugfix updates from Trixie point release
* 13:08 jelto@dns1004: END - running authdns-update
* 13:06 jelto@dns1004: START - running authdns-update
* 13:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 13:04 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:04 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 13:00 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:59 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1066
* 12:59 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1066
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:58 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:54 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 12:53 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2901: Test
* 12:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2901: Test
* 12:52 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'sync'.
* 12:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 12:50 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2901: Test
* 12:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:50 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 12:42 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:41 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-magru and A:durum
* 12:38 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-magru and A:durum
* 12:34 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 12:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1282: Pooling db1282 into s6
* 12:31 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 12:25 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:21 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:19 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:15 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_ulsfo
* 12:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1020.eqiad.wmnet with OS bookworm
* 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: db2207 repool
* 12:07 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_ulsfo
* 12:04 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 11:58 kart_: cxserver: Use urldownloader LVS endpoint ([[phab:T429175|T429175]])
* 11:57 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply
* 11:56 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply
* 11:56 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_magru
* 11:55 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply
* 11:55 moritzm: installing rsync security updates
* 11:55 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply
* 11:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:52 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply
* 11:51 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply
* 11:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1282: Pooling db1282 into s6
* 11:45 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1282 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96328 and previous config saved to /var/cache/conftool/dbconfig/20260903-114526-marostegui.json
* 11:43 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqiad
* 11:35 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqiad
* 11:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1020.eqiad.wmnet with OS bookworm
* 11:26 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 11:24 cgoubert@deploy1003: Finished scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter (duration: 12m 01s)
* 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: db2207 repool
* 11:22 cgoubert@deploy1003: cgoubert: Continuing with deployment
* 11:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqsin
* 11:17 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_esams
* 11:15 cgoubert@deploy1003: cgoubert: mediawiki: enable forward of fatal metrics to statsd exporter synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:14 cgoubert@deploy1003: Started scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter
* 11:10 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_esams
* 11:09 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 11:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1019.eqiad.wmnet with OS bookworm
* 11:01 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_drmrs
* 10:59 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_codfw
* 10:52 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_codfw
* 10:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:41 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 10:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 10:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow2005.codfw.wmnet
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2005.codfw.wmnet with OS trixie
* 10:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:17 btullis@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync
* 10:17 btullis@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync
* 10:16 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1019.eqiad.wmnet with OS bookworm
* 10:15 btullis@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: sync
* 10:15 btullis@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: sync
* 10:12 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s8
* 10:11 marostegui: Move s8 sanitarium from db1167 to db1281 [[phab:T434778|T434778]]
* 10:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 10:03 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 09:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:55 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2003.codfw.wmnet with OS trixie
* 09:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:43 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow2005.codfw.wmnet with OS trixie
* 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:42 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1018.eqiad.wmnet with OS bookworm
* 09:41 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_ulsfo
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:39 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2003.codfw.wmnet with reason: host reimage
* 09:39 hnowlan: fixed currently oncall pane in klaxon
* 09:38 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:38 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434751|T434751]]
* 09:37 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:37 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow2005.codfw.wmnet
* 09:35 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434775|T434775]]
* 09:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s7
* 09:34 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:33 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_ulsfo
* 09:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqiad
* 09:30 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] (duration: 09m 30s)
* 09:27 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:25 zabe@deploy1003: zabe: Continuing with deployment
* 09:25 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqiad
* 09:25 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 09:25 zabe@deploy1003: zabe: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1174 from dbctl [[phab:T436904|T436904]]', diff saved to https://phabricator.wikimedia.org/P96323 and previous config saved to /var/cache/conftool/dbconfig/20260903-092448-marostegui.json
* 09:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:21 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]]
* 09:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqsin
* 09:17 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:15 topranks: put traffic on Lumen codfw<->eqiad link as it is stable [[phab:T435810|T435810]]
* 09:14 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_esams
* 09:09 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:06 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_esams
* 09:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 09:03 marostegui: Move s6 sanitarium from db1165 to db1279 [[phab:T434775|T434775]]
* 09:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1018.eqiad.wmnet with OS bookworm
* 08:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s6
* 08:57 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_drmrs
* 08:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_codfw
* 08:55 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw
* 08:49 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_codfw
* 08:49 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:45 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 08:45 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:42 marostegui: Move s5 sanitarium from db1161 to db1275 [[phab:T434776|T434776]]
* 08:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 24 hosts with reason: Changing sanitarium master in s5
* 08:38 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_magru
* 08:37 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:37 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server (duration: 01m 21s)
* 08:37 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:36 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server
* 08:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:33 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server (duration: 01m 28s)
* 08:32 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server
* 08:27 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 08:11 moritzm: uploaded wmf-laptop 1.0.7 to apt.wikimedia.org
* 08:03 marostegui: Move s2 sanitarium from db1156 to db1271 [[phab:T434287|T434287]]
* 07:59 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2001.codfw.wmnet
* 07:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 23 hosts with reason: Changing sanitarium master in s2
* 07:49 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2001.codfw.wmnet
* 07:39 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:29 chlod: UTC morning backport window done
* 07:27 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] (duration: 11m 54s)
* 07:22 chlod@deploy1003: chlod, tryvix1509: Continuing with deployment
* 07:22 XioNoX: push pfw policies - [[phab:T436729|T436729]]
* 07:20 chlod@deploy1003: chlod, tryvix1509: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]]
* 07:15 marostegui: Power off db1228 for maintenance
* 07:13 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1228.eqiad.wmnet with reason: Onsite maintenance
* 07:01 arnaudb@dns1006: END - running authdns-update
* 06:58 arnaudb@dns1006: START - running authdns-update
* 06:54 jmm@cumin2003: END (PASS) - Cookbook sre.wdqs.restart-nginx-envoy (exit_code=0) rolling restart_daemons on A:wcqs-public
* 06:52 jmm@cumin2003: START - Cookbook sre.wdqs.restart-nginx-envoy rolling restart_daemons on A:wcqs-public
* 06:46 moritzm: installing libxml2 security updates
* 06:27 hashar: Upgrading CI Jenkins on contint1003 # [[phab:T436812|T436812]]
* 06:11 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet
* 06:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts2002.codfw.wmnet
* 06:00 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet
* 05:56 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts2002.codfw.wmnet
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 48s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:16 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1017.eqiad.wmnet with OS bookworm
== 2026-09-02 ==
* 23:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] (duration: 10m 21s)
* 23:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:50 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 23:49 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]]
* 23:38 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:38 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host aqs1017.eqiad.wmnet with OS bookworm
* 22:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] (duration: 11m 03s)
* 22:32 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 22:30 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:26 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]]
* 22:05 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] (duration: 14m 14s)
* 21:59 krinkle@deploy1003: krinkle: Continuing with deployment
* 21:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:50 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]]
* 21:44 inflatador: bking@apt1002 sudo -E private_reprepro --ignore=wrongdistribution -C matomo_plugins include bookworm-wikimedia-private matomo-plugin-customreports_5.5.0-1_amd64.changes [[phab:T431608|T431608]]
* 21:40 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] (duration: 09m 48s)
* 21:35 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:34 jforrester@deploy1003: jforrester: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:34 inflatador: bking@apt1002 sudo -E reprepro -C main include bookworm-wikimedia matomo-plugin-marketingcampaignsreporting_5.2.2-3_amd64.changes [[phab:T431608|T431608]]
* 21:30 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]]
* 21:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 21:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:03 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 20:50 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 20:27 dancy@deploy1003: Finished scap sync-world: testing (duration: 09m 21s)
* 20:18 dancy@deploy1003: Started scap sync-world: testing
* 20:18 dancy@deploy1003: Installation of scap version "4.288.0" completed for 3 hosts
* 20:16 dancy@deploy1003: Installing scap version "4.288.0" for 3 host(s)
* 19:57 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] (duration: 64m 27s)
* 19:55 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 18:57 jforrester@deploy1003: jforrester: Continuing with deployment
* 18:57 jforrester@deploy1003: jforrester: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:55 swfrench-wmf: deleted pods coredns-85b4f68d95-pk5sn coredns-85b4f68d95-22ddb coredns-85b4f68d95-49k5p in eqiad due to intermittent upstream resolution health check failures correlated with high DNS resolution latency
* 18:53 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]]
* 18:31 sukhe@dns1004: END - running authdns-update
* 18:28 sukhe@dns1004: START - running authdns-update
* 18:26 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:26 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:17 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:16 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 18:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:14 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:14 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqiad
* 18:14 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:49 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqiad
* 17:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:45 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqsin
* 17:38 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:38 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:20 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqsin
* 17:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 17:00 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 16:59 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:48 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 16:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:41 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 16:41 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 16:41 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:38 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:34 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 16:29 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_drmrs
* 16:24 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 16:14 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_drmrs
* 16:11 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-unlock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 16:10 root@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]] (duration: 48m 09s)
* 16:10 root@deploy1003: Forcefully removing global lock: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 16:10 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-unlock-scap for datacenter switchover from codfw to eqiad
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:59 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:58 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters for datacenter switchover from codfw to eqiad
* 15:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 15:58 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-restore-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-restore-ttl for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-start-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 root@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-start-maintenance for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-jobrunner: sync
* 15:56 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-jobrunner: sync
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.07-set-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period ends at: 2026-09-02 15:56:13.434320
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.07-set-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.03-set-db-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.03-set-db-readonly for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.02-set-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period starts at: 2026-09-02 15:53:47.690918
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.02-set-readonly for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.01-stop-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.01-stop-maintenance for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-reduce-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:47 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-reduce-ttl for datacenter switchover from codfw to eqiad
* 15:46 slyngshede@cumin1003: END (ERROR) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=97) for datacenter switchover from codfw to eqiad
* 15:45 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 15:44 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 15:42 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service
* 15:39 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:38 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:31 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:28 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:27 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 15:27 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from codfw to eqiad
* 15:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-lock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:22 root@deploy1003: Locking from deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-lock-scap for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks for datacenter switchover from codfw to eqiad
* 15:17 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service
* 15:15 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_magru
* 15:11 moritzm: import jenkins 2.568.3 to thirdparty/jenkins for trixie-wikimedia
* 14:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 14:44 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_magru
* 14:32 moritzm: installing pdns-recursor security updates
* 14:27 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:27 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:20 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:15 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 14:12 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 14:09 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:09 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:09 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:08 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:06 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:44 moritzm: bounce tcpircbot-logmsgbot/tcpircbot-logmsgbot_cloud on alert1002 to allow cumin1004 [[phab:T427897|T427897]]
* 13:36 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] (duration: 09m 52s)
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Continuing with deployment
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:26 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]]
* 13:25 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 13:24 moritzm: installing wireshark security updates
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 13:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:04 moritzm: import librsvg 2.60.0+dfsg-1+wmf13u1 to component/thumbor for trixie-wikimedia [[phab:T436505|T436505]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 12:44 atsuko@dns1004: END - running authdns-update
* 12:41 atsuko@dns1004: START - running authdns-update
* 12:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:35 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] (duration: 12m 50s)
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:24 dreamyjazz@deploy1003: dreamyjazz, btullis: Continuing with deployment
* 12:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:22 dreamyjazz@deploy1003: dreamyjazz, btullis: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:20 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:17 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]]
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 11:31 marostegui@cumin1003: Removing db1172 from zarcillo [[phab:T436763|T436763]]
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1172.eqiad.wmnet
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:30 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:26 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:26 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:25 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 11:25 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:24 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:20 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1172.eqiad.wmnet
* 11:20 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 11:12 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:11 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:10 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:09 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:05 slyngshede@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.*
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 11:05 slyngshede@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 11:03 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply
* 11:03 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
* 10:52 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:49 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:48 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Pool back db1242
* 10:45 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow4003.ulsfo.wmnet with OS trixie
* 10:31 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1172 from dbctl [[phab:T436763|T436763]]', diff saved to https://phabricator.wikimedia.org/P96318 and previous config saved to /var/cache/conftool/dbconfig/20260902-103152-marostegui.json
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 blake@deploy1003: Finished scap sync-world: non-build deployment for [[phab:T417800|T417800]] (duration: 05m 37s)
* 10:06 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:05 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:04 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:03 blake@deploy1003: Started scap sync-world: non-build deployment for [[phab:T417800|T417800]]
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 10:00 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:00 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:59 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:58 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:57 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:56 jmm@dns1004: END - running authdns-update
* 09:56 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:55 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:54 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:53 jmm@dns1004: START - running authdns-update
* 09:51 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:47 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1228 to dbctl [[phab:T435892|T435892]]', diff saved to https://phabricator.wikimedia.org/P96313 and previous config saved to /var/cache/conftool/dbconfig/20260902-094713-marostegui.json
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:42 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow4003.ulsfo.wmnet with OS trixie
* 09:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow7002.magru.wmnet with OS trixie
* 09:31 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:30 moritzm: installing openjdk-21 security updates
* 09:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:22 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:17 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:17 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:16 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:10 moritzm: installing openjdk-8 security updates
* 09:09 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:08 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:46 tappof: bump space for prometheus k8s-dse in eqiad
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:39 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow7002.magru.wmnet with OS trixie
* 08:36 moritzm: installing libgraphite2 security updates
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:23 Msz2001: UTC morning backport window done
* 08:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] (duration: 14m 36s)
* 08:19 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:19 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet
* 08:14 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:14 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet
* 08:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet
* 08:08 fabfur@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cp5022.eqsin.wmnet with reason: investigating
* 08:07 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]]
* 08:07 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp5022.*
* 08:07 fabfur: depooling and silencing cp5022 ([[phab:T414411|T414411]])
* 08:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet
* 08:03 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:03 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* {{safesubst:SAL entry|1=08:03 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)}}
* 07:49 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 07:49 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]] synced to the
* 07:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1004.eqiad.wmnet
* 07:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1004.eqiad.wmnet
* 07:30 jmm@dns1004: END - running authdns-update
* {{safesubst:SAL entry|1=07:27 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]}}
* 07:27 jmm@dns1004: START - running authdns-update
* 07:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] (duration: 16m 04s)
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1242: Cloning db1228
* 07:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1242: Cloning db1228
* 07:18 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1228,1242].eqiad.wmnet with reason: db1242 needs to clone db1228
* 07:17 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Continuing with deployment
* 07:12 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1228.eqiad.wmnet with OS trixie
* 07:10 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be veri
* 07:06 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]]
* 06:43 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: build: Updating npm dependencies (duration: 00m 13s)
* 06:43 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: build: Updating npm dependencies
* 06:39 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:32 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:23 slyngshede@dns1004: END - running authdns-update
* 06:21 marostegui: Drop cu* tables from s3 bswiktionary [[phab:T435965|T435965]]
* 06:20 slyngshede@dns1004: START - running authdns-update
* 06:18 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1228.eqiad.wmnet with OS trixie
* 06:13 XioNoX: re-enable magru cr1/asw1-b3 link - [[phab:T436675|T436675]]
* 05:06 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] (duration: 04m 42s)
* 05:04 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:03 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:01 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 05:01 tstarling@deploy1003: Scap cancelled without rolling back.
* 04:53 tstarling@deploy1003: tstarling: Continuing with deployment
* 04:29 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 04:25 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 43s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-01 ==
* 21:59 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] (duration: 18m 05s)
* 21:52 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:47 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]]
* 21:38 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] (duration: 23m 55s)
* 21:28 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:20 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:14 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]]
* 20:47 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1016.eqiad.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:24 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:11 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 20:11 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:57 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:57 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:56 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:54 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:52 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:47 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:45 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:42 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:41 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:40 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:39 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:32 jhancock@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudcephosd1055
* 19:32 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:30 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.* [reason: update IP addrs]
* 19:30 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 19:30 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:23 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:22 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:16 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:13 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1056
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1056
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1055
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:12 jclark@cumin1003: END (ERROR) - Cookbook sre.network.configure-switch-interfaces (exit_code=97) for host cloudcephosd1054
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1054
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:11 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:06 jclark@cumin1003: START - Cookbook sre.dns.netbox
* 19:06 sukhe@dns1004: END - running authdns-update
* 19:03 sukhe@dns1004: START - running authdns-update
* 18:14 dancy@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 17:04 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 17:00 dancy@deploy1003: Finished scap sync-world: testing (duration: 08m 07s)
* 16:52 dancy@deploy1003: Started scap sync-world: testing
* 16:48 dancy@deploy1003: sync-world aborted: testing (duration: 00m 05s)
* 16:48 dancy@deploy1003: Started scap sync-world: testing
* 16:47 dancy@deploy1003: Installation of scap version "4.287.0" completed for 156 hosts
* 16:42 dancy@deploy1003: Installing scap version "4.287.0" for 156 host(s)
* 16:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 16:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 15:51 moritzm: installing mesa security updates
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts phab1004.eqiad.wmnet
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:27 aokoth@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:20 aokoth@cumin1003: START - Cookbook sre.dns.netbox
* 15:14 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595] (duration: 05m 32s)
* 15:14 aokoth@cumin1003: START - Cookbook sre.hosts.decommission for hosts phab1004.eqiad.wmnet
* 15:11 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on phab1004.eqiad.wmnet with reason: Decom
* 15:09 joal@deploy1003: Started deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595]
* 15:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config: apply
* 15:08 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config: apply
* 14:55 hashar: Restarted Jenkins on releases1003
* 14:51 hashar: Restarted CI Jenkins on contint1003
* 14:48 hashar: Restarting Gerrit primary on gerrit2003
* 14:45 hashar: Restarted Gerrit on gerrit1003 and gerrit2002
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 14:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 14:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 14:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:24 moritzm: installing curl security updates
* 14:24 jmm@dns1004: END - running authdns-update
* 14:23 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser (duration: 00m 14s)
* 14:23 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:21 jmm@dns1004: START - running authdns-update
* 14:21 jmm@dns1004: END - running authdns-update
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 14:19 jmm@dns1004: START - running authdns-update
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet
* 14:14 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:14 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:13 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595] (duration: 07m 26s)
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Test
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Test
* 14:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet
* 14:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Test
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Test
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 14:05 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595]
* 14:05 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] (duration: 37m 37s)
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 hashar: Removed openjdk-17 packages from contint1002/contint2002 following relocation of CI Jenkins to contint1003/contint2003 # [[phab:T418521|T418521]]
* 14:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:02 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:02 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:00 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595] (duration: 00m 45s)
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 13:59 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595]
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:58 ladsgroup@dns1004: END - running authdns-update
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:57 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 ladsgroup@dns1004: START - running authdns-update
* 13:56 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:56 ladsgroup@dns1004: END - running authdns-update
* 13:55 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:53 ladsgroup@dns1004: START - running authdns-update
* 13:50 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:49 kharlan@deploy1003: kharlan: Continuing with deployment
* 13:48 kharlan@deploy1003: kharlan: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2004.wikimedia.org
* 13:42 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:41 jmm@dns1004: END - running authdns-update
* 13:40 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2004.wikimedia.org
* 13:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2003.wikimedia.org
* 13:38 jmm@dns1004: START - running authdns-update
* 13:34 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2003.wikimedia.org
* 13:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1004.wikimedia.org
* 13:26 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]]
* 13:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:25 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:24 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1004.wikimedia.org
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:23 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] (duration: 20m 24s)
* 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1003.wikimedia.org
* 13:20 fnegri@deploy1003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply
* 13:19 fnegri@deploy1003: helmfile [eqiad] START helmfile.d/services/toolhub: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:16 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1003.wikimedia.org
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:15 moritzm: bump urldownloader[12]00[34] to 8G RAM [[phab:T429175|T429175]]
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:15 fnegri@deploy1003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 fnegri@deploy1003: helmfile [codfw] START helmfile.d/services/toolhub: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:11 fnegri@deploy1003: helmfile [staging] DONE helmfile.d/services/toolhub: apply
* 13:11 aude@deploy1003: aude: Continuing with deployment
* 13:10 fnegri@deploy1003: helmfile [staging] START helmfile.d/services/toolhub: apply
* 13:07 aude@deploy1003: aude: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2004.wikimedia.org with OS bookworm
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow1004.eqiad.wmnet
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1004.eqiad.wmnet with OS trixie
* 12:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:47 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:41 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] (duration: 16m 25s)
* 12:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 12:34 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 12:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 12:33 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 12:31 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:29 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:25 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]]
* 12:22 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1004.eqiad.wmnet with OS trixie
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:20 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 12:20 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 12:16 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow1004.eqiad.wmnet
* 12:15 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2004.wikimedia.org with OS bookworm
* 12:14 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 12:07 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2003.wikimedia.org with OS bookworm
* 11:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:43 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:32 jmm@cumin2003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling restart_daemons on A:kafka-test-eqiad
* 11:26 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2003.wikimedia.org with OS bookworm
* 11:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1004.wikimedia.org with OS bookworm
* 11:12 moritzm: installing openjdk-21 security updates
* 11:12 jmm@cumin2003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling restart_daemons on A:kafka-test-eqiad
* 10:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:53 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:49 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2902: Pool back db2902
* 10:45 moritzm: installing Python 3.11 security updates
* 10:37 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1004.wikimedia.org with OS bookworm
* 10:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5022.eqsin.wmnet with OS trixie
* 10:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp5022.eqsin.wmnet on all recursors
* 10:36 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache cp5022.eqsin.wmnet on all recursors
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:30 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:04 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:04 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2902: Pool back db2902
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:03 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2902: test
* 10:03 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 10:01 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 09:58 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:50 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/echoserver: apply
* 09:49 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/echoserver: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Change db1176 and db2230's weight, test-s4 masters, to 0 to mimic the rest of production [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96292 and previous config saved to /var/cache/conftool/dbconfig/20260901-092444-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96291 and previous config saved to /var/cache/conftool/dbconfig/20260901-090233-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96290 and previous config saved to /var/cache/conftool/dbconfig/20260901-090158-marostegui.json
* 09:01 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96289 and previous config saved to /var/cache/conftool/dbconfig/20260901-090121-marostegui.json
* 09:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:56 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:50 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1003.wikimedia.org with OS bookworm
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1074.eqiad.wmnet
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:46 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:43 marostegui@cumin1003: dbctl commit (dc=all): 'Test repool db2902', diff saved to https://phabricator.wikimedia.org/P96288 and previous config saved to /var/cache/conftool/dbconfig/20260901-084317-marostegui.json
* 08:42 marostegui@cumin1003: dbctl commit (dc=all): 'Test depool db2902', diff saved to https://phabricator.wikimedia.org/P96287 and previous config saved to /var/cache/conftool/dbconfig/20260901-084249-marostegui.json
* 08:39 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:37 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2902: test
* 08:36 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96285 and previous config saved to /var/cache/conftool/dbconfig/20260901-083557-marostegui.json
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96284 and previous config saved to /var/cache/conftool/dbconfig/20260901-083527-marostegui.json
* 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96283 and previous config saved to /var/cache/conftool/dbconfig/20260901-083432-marostegui.json
* 08:32 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1074.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1073.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:27 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:24 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 08:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022']
* 08:16 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1073.eqiad.wmnet
* 08:15 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1003.wikimedia.org with OS bookworm
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1072.eqiad.wmnet
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:14 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:10 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:05 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1072.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1067.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:00 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:56 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:54 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022']
* 07:53 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 07:52 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie
* 07:52 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 07:50 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1067.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1066.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:47 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 07:41 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:34 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1066.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1065.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:32 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:27 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:23 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1065.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1075.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:15 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 06:49 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:45 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1075.eqiad.wmnet
* 06:29 moritzm: installing Java 17 security updates
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.15 (duration: 02m 25s)
* 03:50 denisse@deploy1003: Finished deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2 (duration: 00m 19s)
* 03:50 denisse@deploy1003: Started deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2
* 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]] (duration: 37m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:30 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 00:29 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
== Other archives ==
See [[Server Admin Log/Archives]].
<noinclude>
[[Category:SAL]]
[[Category:Operations]]
</noinclude>
aovt0c6otaot71qbqbuyrqf21b9s1lv
2458702
2458701
2026-09-19T16:55:29Z
Stashbot
7414
ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
2458702
wikitext
text/x-wiki
== 2026-09-19 ==
* 16:55 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:11 urbanecm: Attach SHB@commonswiki to the SUL account manually ([[phab:T438591|T438591]], see [[phab:T438591|T438591]]#12341750 for what I did exactly)
* 04:08 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 04:08 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 04:08 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 04:07 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 36s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-18 ==
* 22:41 rzl@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=sessionstore,name=eqiad
* 17:08 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343065{{!}}mw.DesktopArticleTarget: if source education is enabled suppress welcome (T434249)]] (duration: 09m 26s)
* 17:05 Dreamy_Jazz: Created `securepoll_log` on `nlwiki` main DB cluster for [[phab:T434045|T434045]]
* 17:04 kemayo@deploy1003: kemayo: Continuing with deployment
* 17:03 kemayo@deploy1003: kemayo: Backport for [[gerrit:1343065{{!}}mw.DesktopArticleTarget: if source education is enabled suppress welcome (T434249)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:59 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1343065{{!}}mw.DesktopArticleTarget: if source education is enabled suppress welcome (T434249)]]
* 16:49 oblivian@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343127{{!}}ResourceLoader: hotfix for current logspam over the weekend (T438387)]] (duration: 11m 36s)
* 16:42 oblivian@deploy1003: oblivian: Continuing with deployment
* 16:42 oblivian@deploy1003: oblivian: Backport for [[gerrit:1343127{{!}}ResourceLoader: hotfix for current logspam over the weekend (T438387)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:37 oblivian@deploy1003: Started scap sync-world: Backport for [[gerrit:1343127{{!}}ResourceLoader: hotfix for current logspam over the weekend (T438387)]]
* 16:09 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ml-serve1016.eqiad.wmnet with OS trixie
* 14:49 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie
* 13:37 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host registry1004.eqiad.wmnet with OS trixie
* 13:23 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on registry1004.eqiad.wmnet with reason: host reimage
* 13:18 elukey@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on registry1004.eqiad.wmnet with reason: host reimage
* 13:04 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host registry1004.eqiad.wmnet with OS trixie
* 12:14 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:13 jclark@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:30 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 09:30 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 09:22 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on an-redacteddb1001.eqiad.wmnet with reason: cloning
* 09:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 09:19 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 09:18 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 21 hosts with reason: cloning db1270
* 09:18 marostegui: clone db1270:x4 from db1155:x4 lag will appear on x4
* 09:09 brouberol@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 09:08 brouberol@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 08:23 brouberol@dns1004: END - running authdns-update
* 08:21 brouberol@dns1004: START - running authdns-update
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 57s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-17 ==
* 21:04 tsev@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T438395|T438395]]
* 20:55 tsev@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T438395|T438395]]
* 20:47 jhuneidi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342774{{!}}Worklist Promotion test kitchen - Enable flag in production (T434513)]], [[gerrit:1342781{{!}}Exclude returntoapp query from app interception on iOS (T438395)]], [[gerrit:1342798{{!}}Revert "Update wikimania wordmark for 2026"]] (duration: 35m 59s)
* 20:35 jhuneidi@deploy1003: robertsky, jhuneidi, cmelo, tsev: Continuing with deployment
* 20:31 jhuneidi@deploy1003: robertsky, jhuneidi, cmelo, tsev: Backport for [[gerrit:1342774{{!}}Worklist Promotion test kitchen - Enable flag in production (T434513)]], [[gerrit:1342781{{!}}Exclude returntoapp query from app interception on iOS (T438395)]], [[gerrit:1342798{{!}}Revert "Update wikimania wordmark for 2026"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:11 jhuneidi@deploy1003: Started scap sync-world: Backport for [[gerrit:1342774{{!}}Worklist Promotion test kitchen - Enable flag in production (T434513)]], [[gerrit:1342781{{!}}Exclude returntoapp query from app interception on iOS (T438395)]], [[gerrit:1342798{{!}}Revert "Update wikimania wordmark for 2026"]]
* 19:15 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1100.eqiad.wmnet with OS trixie
* 19:15 vriley@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1004"
* 19:10 vriley@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1004"
* 18:52 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1100.eqiad.wmnet with reason: host reimage
* 18:48 vriley@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1100.eqiad.wmnet with reason: host reimage
* 18:32 vriley@cumin1004: START - Cookbook sre.hosts.reimage for host ms-be1100.eqiad.wmnet with OS trixie
* 18:20 urbanecm: Deploy a security fix for [[phab:T438389|T438389]]
* 17:47 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be1100.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:38 vriley@cumin1004: START - Cookbook sre.hosts.provision for host ms-be1100.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:38 vriley@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1100.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:37 vriley@cumin1004: START - Cookbook sre.hosts.provision for host ms-be1100.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:23 vriley@cumin1004: START - Cookbook sre.hosts.reimage for host ms-be1100.eqiad.wmnet with OS trixie
* 16:52 aokoth@deploy1003: Finished deploy [phabricator/deployment@c386249]: Deploy Phab (duration: 00m 12s)
* 16:52 aokoth@deploy1003: Started deploy [phabricator/deployment@c386249]: Deploy Phab
* 16:42 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1099.eqiad.wmnet with OS trixie
* 16:42 vriley@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1004"
* 16:42 vriley@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1004"
* 16:35 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be1100.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:31 aokoth@deploy1003: Finished deploy [phabricator/deployment@c386249]: Deploy Phab (duration: 00m 19s)
* 16:31 aokoth@deploy1003: Started deploy [phabricator/deployment@c386249]: Deploy Phab
* 16:21 vriley@cumin1004: START - Cookbook sre.hosts.provision for host ms-be1100.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:20 vriley@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be1100
* 16:20 vriley@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host ms-be1100
* 16:19 vriley@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:19 vriley@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [ms-be1100] - vriley@cumin1004"
* 16:19 vriley@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [ms-be1100] - vriley@cumin1004"
* 16:15 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 16:15 vriley@cumin1004: START - Cookbook sre.dns.netbox
* 16:15 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 16:14 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 16:14 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 16:12 moritzm: installing libapache-mod-auth-oidc security updates
* 16:12 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1099.eqiad.wmnet with reason: host reimage
* 16:11 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 16:11 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 16:08 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/editcheck-headless: apply
* 16:07 rzl@deploy1003: helmfile [staging] START helmfile.d/services/editcheck-headless: apply
* 16:06 vriley@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1099.eqiad.wmnet with reason: host reimage
* 16:01 moritzm: installing aom security updates
* 16:01 btullis@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ceph-admin2001.codfw.wmnet
* 16:01 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ceph-admin2001.codfw.wmnet with OS bookworm
* 15:55 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 15:55 rzl@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 15:54 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 15:54 rzl@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 15:54 rzl@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 15:53 rzl@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 15:53 rzl@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 15:52 rzl@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 15:51 vriley@cumin1004: START - Cookbook sre.hosts.reimage for host ms-be1099.eqiad.wmnet with OS trixie
* 15:48 moritzm: installing libde265 security updates
* 15:44 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ceph-admin2001.codfw.wmnet with reason: host reimage
* 15:40 btullis@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ceph-admin1001.eqiad.wmnet
* 15:40 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ceph-admin1001.eqiad.wmnet with OS bookworm
* 15:39 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir4003.*
* 15:37 btullis@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ceph-admin2001.codfw.wmnet with reason: host reimage
* 15:26 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir4003.ulsfo.wmnet with OS trixie
* 15:25 aokoth@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host phab2003.codfw.wmnet with OS trixie
* 15:23 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ceph-admin1001.eqiad.wmnet with reason: host reimage
* 15:19 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host registry1005.eqiad.wmnet with OS trixie
* 15:17 btullis@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ceph-admin1001.eqiad.wmnet with reason: host reimage
* 15:16 btullis@cumin1004: START - Cookbook sre.hosts.reimage for host ceph-admin2001.codfw.wmnet with OS bookworm
* 15:16 btullis@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ceph-admin2001.codfw.wmnet - btullis@cumin1004"
* 15:16 btullis@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ceph-admin2001.codfw.wmnet - btullis@cumin1004"
* 15:15 btullis@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ceph-admin2001.codfw.wmnet on all recursors
* 15:15 btullis@cumin1004: START - Cookbook sre.dns.wipe-cache ceph-admin2001.codfw.wmnet on all recursors
* 15:15 btullis@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:15 btullis@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ceph-admin2001.codfw.wmnet - btullis@cumin1004"
* 15:15 btullis@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ceph-admin2001.codfw.wmnet - btullis@cumin1004"
* 15:08 aokoth@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on phab2003.codfw.wmnet with reason: host reimage
* 15:06 Msz2001: Deployed private code changes to Suggestedinvestigations
* 15:05 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir4003.ulsfo.wmnet with reason: host reimage
* 15:05 aokoth@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on phab2003.codfw.wmnet with reason: host reimage
* 15:04 btullis@cumin1004: START - Cookbook sre.hosts.reimage for host ceph-admin1001.eqiad.wmnet with OS bookworm
* 15:04 btullis@cumin1004: START - Cookbook sre.dns.netbox
* 15:04 btullis@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ceph-admin1001.eqiad.wmnet - btullis@cumin1004"
* 15:04 btullis@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ceph-admin1001.eqiad.wmnet - btullis@cumin1004"
* 15:04 btullis@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ceph-admin1001.eqiad.wmnet on all recursors
* 15:04 btullis@cumin1004: START - Cookbook sre.dns.wipe-cache ceph-admin1001.eqiad.wmnet on all recursors
* 15:04 btullis@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:04 btullis@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ceph-admin1001.eqiad.wmnet - btullis@cumin1004"
* 15:04 btullis@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ceph-admin1001.eqiad.wmnet - btullis@cumin1004"
* 15:02 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on registry1005.eqiad.wmnet with reason: host reimage
* 15:01 btullis@cumin1004: START - Cookbook sre.ganeti.makevm for new host ceph-admin2001.codfw.wmnet
* 15:00 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342265{{!}}JsonSchemaBuilder: Cache the root schema in the process (T437588)]], [[gerrit:1342264{{!}}JsonSchemaBuilder: Cache the root schema in the process (T437588)]], [[gerrit:1342684{{!}}SI: Preserve the username filter when switching queues (T438308)]] (duration: 12m 34s)
* 15:00 btullis@cumin1004: START - Cookbook sre.dns.netbox
* 15:00 btullis@cumin1004: START - Cookbook sre.ganeti.makevm for new host ceph-admin1001.eqiad.wmnet
* 14:58 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir4003.ulsfo.wmnet with reason: host reimage
* 14:58 elukey@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on registry1005.eqiad.wmnet with reason: host reimage
* 14:55 urbanecm@deploy1003: mszwarc, urbanecm: Continuing with deployment
* 14:51 urbanecm@deploy1003: mszwarc, urbanecm: Backport for [[gerrit:1342265{{!}}JsonSchemaBuilder: Cache the root schema in the process (T437588)]], [[gerrit:1342264{{!}}JsonSchemaBuilder: Cache the root schema in the process (T437588)]], [[gerrit:1342684{{!}}SI: Preserve the username filter when switching queues (T438308)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:51 aokoth@cumin1004: START - Cookbook sre.hosts.reimage for host phab2003.codfw.wmnet with OS trixie
* 14:50 aokoth@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on phab2003.codfw.wmnet with reason: Reimage
* 14:47 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1342265{{!}}JsonSchemaBuilder: Cache the root schema in the process (T437588)]], [[gerrit:1342264{{!}}JsonSchemaBuilder: Cache the root schema in the process (T437588)]], [[gerrit:1342684{{!}}SI: Preserve the username filter when switching queues (T438308)]]
* 14:42 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1331830{{!}}Keep Balinese Palm Leaf variants enabled on wikisource (T436398)]], [[gerrit:1340216{{!}}Turn on variant conversion for PageAssessments (T328012)]], [[gerrit:1341949{{!}}Parsoid Read Views: Enable on 61 wikiquote wikis (T437917)]] (duration: 15m 31s)
* 14:39 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host registry1005.eqiad.wmnet with OS trixie
* 14:35 cscott@deploy1003: ssastry, cscott: Continuing with deployment
* 14:33 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir4003.ulsfo.wmnet with OS trixie
* 14:32 cscott@deploy1003: ssastry, cscott: Backport for [[gerrit:1331830{{!}}Keep Balinese Palm Leaf variants enabled on wikisource (T436398)]], [[gerrit:1340216{{!}}Turn on variant conversion for PageAssessments (T328012)]], [[gerrit:1341949{{!}}Parsoid Read Views: Enable on 61 wikiquote wikis (T437917)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:26 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1331830{{!}}Keep Balinese Palm Leaf variants enabled on wikisource (T436398)]], [[gerrit:1340216{{!}}Turn on variant conversion for PageAssessments (T328012)]], [[gerrit:1341949{{!}}Parsoid Read Views: Enable on 61 wikiquote wikis (T437917)]]
* 14:20 caro@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342678{{!}}enwiki desktop VE: add education popup for switching to source editor (T434249)]], [[gerrit:1342362{{!}}Make VE the default editor on enwiki desktop (T436574)]] (duration: 33m 52s)
* 14:07 caro@deploy1003: caro: Continuing with deployment
* 14:06 caro@deploy1003: caro: Backport for [[gerrit:1342678{{!}}enwiki desktop VE: add education popup for switching to source editor (T434249)]], [[gerrit:1342362{{!}}Make VE the default editor on enwiki desktop (T436574)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:46 caro@deploy1003: Started scap sync-world: Backport for [[gerrit:1342678{{!}}enwiki desktop VE: add education popup for switching to source editor (T434249)]], [[gerrit:1342362{{!}}Make VE the default editor on enwiki desktop (T436574)]]
* 13:37 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 13:34 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:33 elukey@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:31 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:30 elukey@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:29 Emperor: apus - radosgw-admin quota set --quota-scope=user --uid=docker-registry --max-size=5T [[phab:T438339|T438339]]
* 13:24 jclark@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1016
* 13:24 jclark@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1016
* 13:11 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:11 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:09 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:09 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:55 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:53 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 12:42 jclark@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:40 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:40 jclark@cumin1003: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:35 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 12:19 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342626{{!}}MathMathML: Simplify Mathoid fallback/a11y class logic (T436026)]], [[gerrit:1342627{{!}}ext.math.mathjax: Implement mwe-math-mathml-a11y for client-side MathJax (T436026)]] (duration: 13m 04s)
* 12:14 krinkle@deploy1003: krinkle: Continuing with deployment
* 12:10 krinkle@deploy1003: krinkle: Backport for [[gerrit:1342626{{!}}MathMathML: Simplify Mathoid fallback/a11y class logic (T436026)]], [[gerrit:1342627{{!}}ext.math.mathjax: Implement mwe-math-mathml-a11y for client-side MathJax (T436026)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:05 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1342626{{!}}MathMathML: Simplify Mathoid fallback/a11y class logic (T436026)]], [[gerrit:1342627{{!}}ext.math.mathjax: Implement mwe-math-mathml-a11y for client-side MathJax (T436026)]]
* 10:37 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 10:28 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 10:10 blake@deploy1003: Finished scap sync-world: cleanup for [[phab:T417800|T417800]] (duration: 03m 57s)
* 10:07 blake@deploy1003: Started scap sync-world: cleanup for [[phab:T417800|T417800]]
* 09:52 marostegui@cumin1004: dbctl commit (dc=all): 'Fix weights [[phab:T436496|T436496]]', diff saved to https://phabricator.wikimedia.org/P96466 and previous config saved to /var/cache/conftool/dbconfig/20260917-095235-marostegui.json
* 09:51 marostegui@cumin1004: dbctl commit (dc=all): 'Fix weights [[phab:T436496|T436496]]', diff saved to https://phabricator.wikimedia.org/P96465 and previous config saved to /var/cache/conftool/dbconfig/20260917-095131-marostegui.json
* 09:41 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:41 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 09:41 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:40 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 09:40 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:40 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 09:35 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 09:35 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 08:37 moritzm: pruned obsolete Bullseye image prometheus-nutcracker-exporter from the docker registry [[phab:T416452|T416452]]
* 08:34 XioNoX: Manually install gnmic 0.49.0 on netflow2005 - [[phab:T438291|T438291]]
* 08:28 brouberol@dns1004: END - running authdns-update
* 08:26 brouberol@dns1004: START - running authdns-update
* 08:13 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.20 refs [[phab:T430839|T430839]]
* 08:10 moritzm: imported nodejs_26.8.2-1nodesource1 to thirdparty/node26 for trixie-wikimedia [[phab:T437510|T437510]]
* 08:07 Amir1: dropped links tables from db2206 ([[phab:T437278|T437278]])
* 08:03 Amir1: dropped links tables from db2219 ([[phab:T437278|T437278]])
* 08:01 Amir1: dropped links tables from db2236 ([[phab:T437278|T437278]])
* 07:59 Amir1: dropped non-links tables from db1262 ([[phab:T437278|T437278]])
* 07:57 Amir1: dropped non-links tables from db2245 ([[phab:T437278|T437278]])
* 07:52 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@8eaca67] (releasing): [[phab:T438205|T438205]] to prod host (duration: 00m 44s)
* 07:52 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@8eaca67] (releasing): [[phab:T438205|T438205]] to prod host
* 07:49 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@8eaca67] (releasing): [[phab:T438205|T438205]] to backup host (duration: 00m 47s)
* 07:48 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@8eaca67] (releasing): [[phab:T438205|T438205]] to backup host
* 07:25 XioNoX: Manually install gnmic 0.49.0 on netflow1004 - [[phab:T438291|T438291]]
* 07:24 mlitn@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342398{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342396{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342395{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342545{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342546{{!}}Localisation updates from https://translatewiki.net.]],
* 07:19 mlitn@deploy1003: mlitn, jdlrobson: Continuing with deployment
* {{safesubst:SAL entry|1=07:18 mlitn@deploy1003: mlitn, jdlrobson: Backport for [[gerrit:1342398{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342396{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342395{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342545{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342546{{!}}Localisation updates from https://translatewiki.net.]], [[gerri}}
* 07:11 mlitn@deploy1003: Started scap sync-world: Backport for [[gerrit:1342398{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342396{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342395{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342545{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342546{{!}}Localisation updates from https://translatewiki.net.]],
* 07:10 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Jmoore111 out of all services on: 2444 hosts
* 06:07 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 05:53 marostegui@cumin1004: END (FAIL) - Cookbook sre.mysql.decommission (exit_code=99)
* 05:53 marostegui@cumin1004: Removing db1180 from zarcillo [[phab:T437222|T437222]]
* 05:53 marostegui@cumin1004: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1180.eqiad.wmnet
* 05:53 marostegui@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 05:53 marostegui@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1180.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1004"
* 05:53 marostegui@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1180.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1004"
* 05:49 marostegui@cumin1004: START - Cookbook sre.dns.netbox
* 05:44 marostegui@cumin1004: START - Cookbook sre.hosts.decommission for hosts db1180.eqiad.wmnet
* 05:43 marostegui@cumin1004: START - Cookbook sre.mysql.decommission
* 04:26 aokoth@cumin1004: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet
* 04:24 aokoth@cumin1004: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet
== 2026-09-16 ==
* 23:10 rzl: rzl@deploy1003 Finished scap sync-world: Backport for [[gerrit:1342091{{!}}Repool poolcounter[1007,2006] (T435163)]] (duration: 11m 09s)
* 22:50 rzl@deploy1003: Started scap sync-world: Backport for [[gerrit:1342091{{!}}Repool poolcounter[1007,2006] (T435163)]]
* 22:47 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342383{{!}}DonorIdentification: Confirm before unlinking donor status in preferences (T436698)]], [[gerrit:1342385{{!}}Make learn more link to new window (T438252)]] (duration: 35m 21s)
* 22:35 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 22:33 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1342383{{!}}DonorIdentification: Confirm before unlinking donor status in preferences (T436698)]], [[gerrit:1342385{{!}}Make learn more link to new window (T438252)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:12 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1342383{{!}}DonorIdentification: Confirm before unlinking donor status in preferences (T436698)]], [[gerrit:1342385{{!}}Make learn more link to new window (T438252)]]
* 22:10 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2006.codfw.wmnet
* 22:06 rzl@cumin2003: START - Cookbook sre.hosts.reboot-single for host poolcounter2006.codfw.wmnet
* 22:06 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1007.eqiad.wmnet
* 22:02 rzl@cumin2003: START - Cookbook sre.hosts.reboot-single for host poolcounter1007.eqiad.wmnet
* 21:56 rzl@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342090{{!}}Repool poolcounter[1006,2005]; depool poolcounter[1007,2006] for reboot (T435163)]] (duration: 09m 39s)
* 21:52 rzl@deploy1003: rzl: Continuing with deployment
* 21:51 rzl@deploy1003: rzl: Backport for [[gerrit:1342090{{!}}Repool poolcounter[1006,2005]; depool poolcounter[1007,2006] for reboot (T435163)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:47 rzl@deploy1003: Started scap sync-world: Backport for [[gerrit:1342090{{!}}Repool poolcounter[1006,2005]; depool poolcounter[1007,2006] for reboot (T435163)]]
* 21:46 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:43 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2005.codfw.wmnet
* 21:42 vriley@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1099.eqiad.wmnet with OS trixie
* 21:41 vriley@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1098.eqiad.wmnet with OS trixie
* 21:41 vriley@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin2003"
* 21:40 vriley@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin2003"
* 21:40 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:40 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:39 rzl@cumin2003: START - Cookbook sre.hosts.reboot-single for host poolcounter2005.codfw.wmnet
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1006.eqiad.wmnet
* 21:38 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:38 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:37 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:35 rzl@cumin2003: START - Cookbook sre.hosts.reboot-single for host poolcounter1006.eqiad.wmnet
* 21:32 rzl@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342089{{!}}Depool poolcounter[1006,2005] for reboot (T435163)]] (duration: 13m 53s)
* 21:26 rzl@deploy1003: rzl: Continuing with deployment
* 21:25 rzl@deploy1003: rzl: Backport for [[gerrit:1342089{{!}}Depool poolcounter[1006,2005] for reboot (T435163)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:23 vriley@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1098.eqiad.wmnet with reason: host reimage
* 21:18 rzl@deploy1003: Started scap sync-world: Backport for [[gerrit:1342089{{!}}Depool poolcounter[1006,2005] for reboot (T435163)]]
* 21:17 vriley@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1098.eqiad.wmnet with reason: host reimage
* 21:10 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:09 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:09 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:09 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:08 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:02 vriley@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS trixie
* 20:49 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342339{{!}}Reapply "Tell VisualEditor about the app web edit tags", modified]] (duration: 35m 55s)
* 20:37 kemayo@deploy1003: cklimas, kemayo: Continuing with deployment
* 20:33 kemayo@deploy1003: cklimas, kemayo: Backport for [[gerrit:1342339{{!}}Reapply "Tell VisualEditor about the app web edit tags", modified]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:13 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1342339{{!}}Reapply "Tell VisualEditor about the app web edit tags", modified]]
* 19:20 dwisehaupt@dns1005: END - running authdns-update
* 19:18 dwisehaupt@dns1005: START - running authdns-update
* 19:06 vriley@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:57 dwisehaupt@dns1005: END - running authdns-update
* 18:55 vriley@cumin2003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:55 dwisehaupt@dns1005: START - running authdns-update
* 18:44 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 18:42 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 18:37 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 18:35 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 18:26 robh@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:22 robh@cumin2003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:21 dzahn@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 18:21 robh@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:21 robh@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:20 dzahn@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 18:20 dzahn@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 18:18 dzahn@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 18:18 mutante: k8s/miscweb: admin_ng deploy: creating namespace for attribution.wikimedia.org [[phab:T437635|T437635]]
* 18:17 dzahn@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 18:17 dzahn@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 18:17 dzahn@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 18:16 dzahn@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 18:13 cdanis@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "fix known-client creation - cdanis@cumin1003"
* 18:13 cdanis@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: fix known-client creation - cdanis@cumin1003
* 18:12 cdanis@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: fix known-client creation - cdanis@cumin1003
* 18:12 cdanis@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "fix known-client creation - cdanis@cumin1003"
* 18:04 vriley@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:51 vriley@cumin2003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:46 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 17:46 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 17:45 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:45 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:44 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:44 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:35 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1006.eqiad.wmnet with OS trixie
* 17:35 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1003"
* 17:30 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:29 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1003"
* 17:27 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:27 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:23 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:23 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:20 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 17:20 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 17:14 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:13 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1006.eqiad.wmnet with reason: host reimage
* 17:10 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:10 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:10 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:10 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:09 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:07 jclark@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1006.eqiad.wmnet with reason: host reimage
* 17:07 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 17:05 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:05 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:52 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie
* 16:46 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir7003.magru.wmnet
* 16:44 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir7003
* 16:19 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie
* 15:42 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply
* 15:42 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply
* 15:42 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 15:42 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir7003.magru.wmnet with OS trixie
* 15:41 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 15:36 jnuche@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342279{{!}}Use parser output value instead of status (T438154)]] (duration: 33m 21s)
* 15:35 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply
* 15:35 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply
* 15:24 jnuche@deploy1003: jnuche, jforrester: Continuing with deployment
* 15:23 jnuche@deploy1003: jnuche, jforrester: Backport for [[gerrit:1342279{{!}}Use parser output value instead of status (T438154)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:19 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie
* 15:18 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir7003.magru.wmnet with reason: host reimage
* 15:14 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir7003.magru.wmnet with reason: host reimage
* 15:03 jnuche@deploy1003: Started scap sync-world: Backport for [[gerrit:1342279{{!}}Use parser output value instead of status (T438154)]]
* 14:50 moritzm: installing apache2 security updates
* 14:49 slyngshede@cumin1003: conftool action : set/pooled=yes; selector: name=cp5026.eqsin.wmnet
* 14:47 slyngshede@cumin1003: conftool action : set/weight=1; selector: name=cp5026.eqsin.wmnet
* 14:45 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5026.eqsin.wmnet with OS trixie
* 14:44 moritzm: installing python-filelock security updates
* 14:42 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir7003.magru.wmnet with OS trixie
* 14:35 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:35 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:34 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:33 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp6002.drmrs.wmnet
* 14:32 sukhe@puppetserver1001: conftool action : set/weight=100; selector: name=cp6002.drmrs.wmnet,service=ats-be
* 14:32 sukhe@puppetserver1001: conftool action : set/weight=1; selector: name=cp6002.drmrs.wmnet,service=cdn
* 14:29 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6002.drmrs.wmnet with OS trixie
* 14:24 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: sync
* 14:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:24 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:24 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: sync
* 14:14 jforrester@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.19,1.47.0-wmf.20,next --multiversion-image-basename docker-registry.discovery.wmnet/restricte
* 14:14 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:14 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:13 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1342008{{!}}abstractwiki: Add three new articles per community advice to show off the feature (T434227)]]
* 14:10 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: sync
* 14:10 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: sync
* 14:10 jforrester@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.19,1.47.0-wmf.20,next --multiversion-image-basename docker-registry.discovery.wmnet/restricte
* 14:10 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: sync
* 14:10 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: sync
* 14:09 Amir1: dropped links tables on db2237 ([[phab:T437278|T437278]])
* 14:08 Amir1: dropped links tables on db1238 ([[phab:T437278|T437278]])
* 14:07 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1342008{{!}}abstractwiki: Add three new articles per community advice to show off the feature (T434227)]]
* 14:03 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5026.eqsin.wmnet with reason: host reimage
* 14:02 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:02 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:02 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 14:01 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:00 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 13:59 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6002.drmrs.wmnet with reason: host reimage
* 13:56 slyngshede@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5026.eqsin.wmnet with reason: host reimage
* 13:54 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cp6002.drmrs.wmnet with reason: host reimage
* 13:53 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 13:52 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 13:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dse-k8s-etcd[1001-1003].eqiad.wmnet
* 13:48 bking@cumin2003: START - Cookbook sre.hosts.remove-downtime for dse-k8s-etcd[1001-1003].eqiad.wmnet
* 13:46 bking@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM dse-k8s-etcd1001.eqiad.wmnet
* 13:46 bking@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM dse-k8s-etcd1001.eqiad.wmnet
* 13:45 bking@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM dse-k8s-etcd1002.eqiad.wmnet
* 13:41 bking@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM dse-k8s-etcd1002.eqiad.wmnet
* 13:41 bking@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM dse-k8s-etcd1003.eqiad.wmnet
* 13:38 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp6002.drmrs.wmnet with OS trixie
* 13:37 bking@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM dse-k8s-etcd1003.eqiad.wmnet
* 13:37 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.reboot-vm (exit_code=99) for VM dse-k8s-etcd1003.eqiad.wmnet
* 13:37 bking@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM dse-k8s-etcd1003.eqiad.wmnet
* 13:34 slyngshede@cumin1003: START - Cookbook sre.hosts.reimage for host cp5026.eqsin.wmnet with OS trixie
* 13:34 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 13:33 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5026.mgmt.eqsin.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:29 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp6002.mgmt.drmrs.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:25 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on dse-k8s-etcd[1001-1003].eqiad.wmnet with reason: Maintenance to increase vCPUS [[phab:T438084|T438084]]
* 13:24 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:24 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:22 slyngshede@cumin1003: START - Cookbook sre.hosts.provision for host cp5026.mgmt.eqsin.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342231{{!}}SI: Unset all filters on links to cases (T434530)]], [[gerrit:1342234{{!}}SI: Unset all filters on links to cases (T434530)]] (duration: 13m 10s)
* 13:19 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org
* 13:19 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org
* 13:19 sukhe@cumin1004: START - Cookbook sre.hosts.provision for host cp6002.mgmt.drmrs.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:18 stran@deploy1003: stran: Continuing with deployment
* 13:15 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: apply
* 13:15 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: apply
* 13:13 stran@deploy1003: stran: Backport for [[gerrit:1342231{{!}}SI: Unset all filters on links to cases (T434530)]], [[gerrit:1342234{{!}}SI: Unset all filters on links to cases (T434530)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:11 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp5026.eqsin.wmnet
* 13:09 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1342231{{!}}SI: Unset all filters on links to cases (T434530)]], [[gerrit:1342234{{!}}SI: Unset all filters on links to cases (T434530)]]
* 13:09 sukhe@cumin1004: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp6002.drmrs.wmnet
* 13:05 sukhe@cumin1004: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp6002.drmrs.wmnet
* 13:05 sukhe@cumin1004: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts cp6002.drmrs.wmnet
* 13:00 dkertesz@cumin1004: conftool action : set/pooled=yes; selector: name=cp5025.eqsin.wmnet
* 12:59 dkertesz@cumin1004: conftool action : set/weight=1; selector: name=cp5025.eqsin.wmnet
* 12:52 sukhe@cumin1004: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp6002.drmrs.wmnet
* 12:52 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cp6002.mgmt.drmrs.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:51 dkertesz: eqsin pooled again ([[phab:T438052|T438052]])
* 12:49 dkertesz@cumin1004: conftool action : set/pooled=yes; selector: cluster=dnsbox,dc=eqsin
* 12:47 dkertesz@dns1004: END - running authdns-update
* 12:45 dkertesz@dns1004: START - running authdns-update
* 12:43 dkertesz@cumin1004: conftool action : set/pooled=yes; selector: cluster=dnsbox,dc=eqsin,service=authdns-update
* 12:41 dkertesz@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqsin [reason: no reason specified, [[phab:T438052|T438052]]]
* 12:41 dkertesz@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool eqsin [reason: no reason specified, [[phab:T438052|T438052]]]
* 12:38 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a1-eqiad
* 12:38 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a1-eqiad
* 12:34 sukhe@cumin1004: START - Cookbook sre.hosts.provision for host cp6002.mgmt.drmrs.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:34 sukhe@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cp6002.drmrs.wmnet with reason: reimage
* 12:33 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp6002.drmrs.wmnet
* 12:13 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5025.eqsin.wmnet with OS trixie
* 12:12 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 12:11 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-a1-eqiad
* 12:09 cmooney@cumin1004: START - Cookbook sre.network.tls for network device ssw1-a1-eqiad
* 12:01 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.20 refs [[phab:T430839|T430839]]
* 11:59 moritzm: pruned obsolete Bullseye image python3-bullseye from the docker registry [[phab:T416452|T416452]]
* 11:50 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341284{{!}}IS/IS-labs: Set wmgUseModeratorToolkit default false (T431000)]] (duration: 10m 32s)
* 11:46 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ml-lab1002.eqiad.wmnet
* 11:45 samtar@deploy1003: samtar: Continuing with deployment
* 11:44 samtar@deploy1003: samtar: Backport for [[gerrit:1341284{{!}}IS/IS-labs: Set wmgUseModeratorToolkit default false (T431000)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:41 klausman@cumin1003: START - Cookbook sre.hosts.reboot-single for host ml-lab1002.eqiad.wmnet
* 11:39 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1341284{{!}}IS/IS-labs: Set wmgUseModeratorToolkit default false (T431000)]]
* 11:39 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5025.eqsin.wmnet with reason: host reimage
* 11:35 slyngshede@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5025.eqsin.wmnet with reason: host reimage
* 11:34 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:33 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:31 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:31 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:27 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:27 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:26 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:25 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:24 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.20 refs [[phab:T430839|T430839]]
* 11:21 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply
* 11:21 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply
* 11:20 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply
* 11:20 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply
* 11:19 moritzm: kicked off a new run of production-images-weekly-rebuild.service on build2004 (previously some leftovers of buster in the config prevented a complete run)
* 11:17 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
* 11:16 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
* 11:11 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
* 11:11 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
* 11:10 jnuche@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342210{{!}}Revert "Tell VisualEditor about the app web edit tags" (T437736 T438125)]] (duration: 33m 14s)
* 11:10 slyngshede@cumin1003: START - Cookbook sre.hosts.reimage for host cp5025.eqsin.wmnet with OS trixie
* 11:05 marostegui@cumin1004: dbctl commit (dc=all): 'Remove db1180 from dbctl [[phab:T437222|T437222]]', diff saved to https://phabricator.wikimedia.org/P96459 and previous config saved to /var/cache/conftool/dbconfig/20260916-110502-marostegui.json
* 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 11:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:01 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 11:01 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 10:57 jnuche@deploy1003: jnuche: Continuing with deployment
* 10:57 jnuche@deploy1003: jnuche: Backport for [[gerrit:1342210{{!}}Revert "Tell VisualEditor about the app web edit tags" (T437736 T438125)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:37 jnuche@deploy1003: Started scap sync-world: Backport for [[gerrit:1342210{{!}}Revert "Tell VisualEditor about the app web edit tags" (T437736 T438125)]]
* 10:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5025.mgmt.eqsin.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:10 slyngshede@cumin1003: START - Cookbook sre.hosts.provision for host cp5025.mgmt.eqsin.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:02 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 10:02 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 09:58 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 09:49 slyngshede@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on cp5025.eqsin.wmnet with reason: reimaging
* 09:48 slyngshede@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cp5025.eqsin.wmnet with reason: reimaging
* 09:41 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-timeline: apply
* 09:41 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-timeline: apply
* 09:38 moritzm: imported routinator 0.15.2-1trixie to thirdparty/routinator [[phab:T438122|T438122]]
* 09:30 slyngshede@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cp5025.mgmt.eqsin.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:29 slyngshede@cumin1003: START - Cookbook sre.hosts.provision for host cp5025.mgmt.eqsin.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:19 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 09:19 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp5025.eqsin.wmnet
* 09:19 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'sync'.
* 09:18 slyngshede@cumin1003: conftool action : set/pooled=yes; selector: name=cp3074.esams.wmnet
* 09:18 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp3074.esams.wmnet
* 09:15 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 09:12 elukey@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'sync'.
* 09:12 elukey@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'sync'.
* 09:11 elukey@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'sync'.
* 09:11 elukey@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'sync'.
* 09:10 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 09:10 elukey@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 09:09 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 09:09 elukey@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 08:55 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 08:54 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 08:40 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:40 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:36 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqsin [reason: depooling for maintainance, [[phab:T438052|T438052]]]
* 08:36 slyngshede@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool eqsin [reason: depooling for maintainance, [[phab:T438052|T438052]]]
* 08:35 slyngshede@cumin1003: END (FAIL) - Cookbook sre.dns.admin (exit_code=99) DNS admin: depool eqsin [reason: no reason specified, no task ID specified]
* 08:35 slyngshede@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool eqsin [reason: no reason specified, no task ID specified]
* 08:35 slyngshede@cumin1003: conftool action : set/pooled=no; selector: cluster=dnsbox,dc=eqsin
* 08:34 fabfur: start depooling eqsin ([[phab:T438052|T438052]])
* 08:24 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.20 refs [[phab:T430839|T430839]]
* 08:22 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:22 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:14 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.20 refs [[phab:T430839|T430839]]
* 08:11 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 08:11 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 08:11 jnuche@deploy1003: Rolling back deployment
* 08:10 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 08:07 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 07:59 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 07:59 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 07:58 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 07:54 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 07:34 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply
* 07:34 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply
* 07:20 mlitn@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342112{{!}}Adds an instrument for pre-image-carousel-retest (T437076)]], [[gerrit:1342113{{!}}Adds an instrument for pre-image-carousel-retest (T437076)]], [[gerrit:1342117{{!}}Set up instrument for 5-arm test (T437076)]], [[gerrit:1342118{{!}}Set up instrument for 5-arm test (T437076)]] (duration: 10m 56s)
* 07:16 mlitn@deploy1003: mlitn: Continuing with deployment
* 07:15 mlitn@deploy1003: mlitn: Backport for [[gerrit:1342112{{!}}Adds an instrument for pre-image-carousel-retest (T437076)]], [[gerrit:1342113{{!}}Adds an instrument for pre-image-carousel-retest (T437076)]], [[gerrit:1342117{{!}}Set up instrument for 5-arm test (T437076)]], [[gerrit:1342118{{!}}Set up instrument for 5-arm test (T437076)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be veri
* 07:09 mlitn@deploy1003: Started scap sync-world: Backport for [[gerrit:1342112{{!}}Adds an instrument for pre-image-carousel-retest (T437076)]], [[gerrit:1342113{{!}}Adds an instrument for pre-image-carousel-retest (T437076)]], [[gerrit:1342117{{!}}Set up instrument for 5-arm test (T437076)]], [[gerrit:1342118{{!}}Set up instrument for 5-arm test (T437076)]]
* 06:50 moritzm: installing sudo security updates
* 06:47 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply
* 06:37 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply
* 05:12 moritzm: pruned obsolete Bullseye image buildkitd from the docker registry [[phab:T416452|T416452]]
* 04:56 kart_: Updated Apertium to 2026-09-15-084320-production ([[phab:T437213|T437213]])
* 04:54 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/apertium: apply
* 04:54 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/apertium: apply
* 04:50 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/apertium: apply
* 04:49 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/apertium: apply
* 04:45 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/apertium: apply
* 04:45 kartik@deploy1003: helmfile [staging] START helmfile.d/services/apertium: apply
* 04:24 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 04:24 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 04:22 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 04:22 oblivian@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 36s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-15 ==
* 23:09 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply
* 23:08 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply
* 23:08 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply
* 23:08 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply
* 23:07 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 23:07 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 23:07 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 23:07 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 23:06 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 23:06 rzl@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 22:57 sukhe@puppetserver1001: conftool action : set/weight=1; selector: name=cp6001.drmrs.wmnet,service=cdn
* 22:50 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc-wf1002.eqiad.wmnet with OS trixie
* 22:46 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs ([[phab:T436363|T436363]])
* 22:46 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs6003.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:46 brett@cumin2003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs6003.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:46 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs6003.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:46 brett@cumin2003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs6003.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:45 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs6002.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:45 brett@cumin2003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs6002.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:45 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs6002.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:44 brett@cumin2003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs6002.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:44 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs6001.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:44 brett@cumin2003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs6001.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:43 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs6001.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:43 brett@cumin2003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs6001.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:43 brett@cumin2003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs ([[phab:T436363|T436363]])
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc-wf1002.eqiad.wmnet with reason: host reimage
* 22:33 brett@puppetserver1001: conftool action : set/weight=100; selector: name=cp6001.*
* 22:32 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp6001.*
* 22:26 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc-wf1002.eqiad.wmnet with reason: host reimage
* 22:07 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc-wf1002
* 22:07 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-wf1002
* 22:07 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host mc-wf1002
* 22:07 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc-wf1002.eqiad.wmnet 142.48.64.10.in-addr.arpa 2.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:07 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache mc-wf1002.eqiad.wmnet 142.48.64.10.in-addr.arpa 2.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:07 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:07 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc-wf1002 - rzl@cumin2003"
* 22:07 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc-wf1002 - rzl@cumin2003"
* 22:02 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:01 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host mc-wf1002
* 22:01 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host mc-wf1002.eqiad.wmnet with OS trixie
* 21:57 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6001.drmrs.wmnet with OS trixie
* 21:55 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply
* 21:55 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply
* 21:53 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply
* 21:53 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply
* 21:53 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:53 rzl@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:52 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:52 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:48 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:48 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:34 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6001.drmrs.wmnet with reason: host reimage
* 21:30 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp6001.drmrs.wmnet with reason: host reimage
* 21:20 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342049{{!}}MobileFrontend: Add app icons (T434258)]] (duration: 11m 47s)
* 21:15 jdlrobson@deploy1003: jdlrobson, cklimas: Continuing with deployment
* 21:13 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp6001.drmrs.wmnet with OS trixie
* 21:12 jdlrobson@deploy1003: jdlrobson, cklimas: Backport for [[gerrit:1342049{{!}}MobileFrontend: Add app icons (T434258)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:12 brett@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp6001.mgmt.drmrs.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:08 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1342049{{!}}MobileFrontend: Add app icons (T434258)]]
* 20:51 cdobbins@puppetserver1001: conftool action : set/weight=1; selector: name=cp2046.codfw.wmnet
* 20:51 cdobbins@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp2046.codfw.wmnet
* 20:48 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341271{{!}}Parsoid Read Views: Enable on all namespaces on wikitech (labswiki) (T437916)]] (duration: 09m 11s)
* 20:47 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp2046.codfw.wmnet with OS trixie
* 20:43 arlolra@deploy1003: ssastry, arlolra: Continuing with deployment
* 20:42 arlolra@deploy1003: ssastry, arlolra: Backport for [[gerrit:1341271{{!}}Parsoid Read Views: Enable on all namespaces on wikitech (labswiki) (T437916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:38 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1341271{{!}}Parsoid Read Views: Enable on all namespaces on wikitech (labswiki) (T437916)]]
* 20:34 brett@cumin2003: START - Cookbook sre.hosts.provision for host cp6001.mgmt.drmrs.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:30 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp6001.*
* 20:24 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp2046.codfw.wmnet with reason: host reimage
* 20:23 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339741{{!}}Enable ReaderExperiments in eswiki, jawiki, and ptwiki (T438009)]] (duration: 15m 58s)
* 20:20 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cp2046.codfw.wmnet with reason: host reimage
* 20:19 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T436363|T436363]])
* 20:19 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs4010.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:19 brett@cumin2003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs4010.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:19 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs4010.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:19 brett@cumin2003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs4010.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:18 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs4009.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:18 arlolra@deploy1003: lwatson, arlolra: Continuing with deployment
* 20:18 brett@cumin2003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs4009.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:17 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs4009.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:17 brett@cumin2003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs4009.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:17 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs4008.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:17 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp3075.esams.wmnet
* 20:17 sukhe@puppetserver1001: conftool action : set/weight=1; selector: name=cp3075.esams.wmnet
* 20:17 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp1103.eqiad.wmnet
* 20:17 brett@cumin2003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs4008.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:16 sukhe@puppetserver1001: conftool action : set/weight=1; selector: name=cp1103.eqiad.wmnet
* 20:16 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs4008.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:16 brett@cumin2003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs4008.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:16 brett@cumin2003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T436363|T436363]])
* 20:11 arlolra@deploy1003: lwatson, arlolra: Backport for [[gerrit:1339741{{!}}Enable ReaderExperiments in eswiki, jawiki, and ptwiki (T438009)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:10 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) config_reloading A:liberica-ulsfo ([[phab:T436363|T436363]])
* 20:08 brett@cumin2003: START - Cookbook sre.loadbalancer.admin config_reloading A:liberica-ulsfo ([[phab:T436363|T436363]])
* 20:08 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp1103.eqiad.wmnet with OS trixie
* 20:07 inflatador: bking@ganeti1046 sudo gnt-instance modify -B memory=4g,vcpus=4 dse-k8s-etcd100[1-3].eqiad.wmnet [[phab:T438084|T438084]]
* 20:07 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1339741{{!}}Enable ReaderExperiments in eswiki, jawiki, and ptwiki (T438009)]]
* 20:06 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp3075.esams.wmnet with OS trixie
* 20:04 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp7009.*
* 20:04 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host cp2046.codfw.wmnet with OS trixie
* 20:02 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp7009.*
* 20:02 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp7009.*
* 20:02 brett@puppetserver1001: conftool action : set/weight=1; selector: name=cp7009.*
* 20:01 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp7009.magru.wmnet with OS trixie
* 19:53 brett@puppetserver1001: conftool action : set/weight=1; selector: name=cp4045.*
* 19:53 brett@puppetserver1001: conftool action : set/weight=1; selector: name=cp4046.*
* 19:52 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4046.*
* 19:51 cdobbins@puppetserver1001: conftool action : set/pooled=no; selector: name=cp2046.codfw.wmnet
* 19:51 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp4046.ulsfo.wmnet with OS trixie
* 19:49 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4045.*
* 19:45 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp1103.eqiad.wmnet with reason: host reimage
* 19:43 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp4045.ulsfo.wmnet with OS trixie
* 19:41 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp3075.esams.wmnet with reason: host reimage
* 19:39 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cp1103.eqiad.wmnet with reason: host reimage
* 19:37 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp7009.magru.wmnet with reason: host reimage
* 19:33 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cp3075.esams.wmnet with reason: host reimage
* 19:32 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp7009.magru.wmnet with reason: host reimage
* 19:27 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp4046.ulsfo.wmnet with reason: host reimage
* 19:23 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp4046.ulsfo.wmnet with reason: host reimage
* 19:21 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp1103.eqiad.wmnet with OS trixie
* 19:19 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp1103.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:19 sukhe@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cp1103.eqiad.wmnet with reason: reimage
* 19:18 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp4045.ulsfo.wmnet with reason: host reimage
* 19:13 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 19:12 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 19:12 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 19:12 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp3075.esams.wmnet with OS trixie
* 19:12 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp4045.ulsfo.wmnet with reason: host reimage
* 19:12 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 19:11 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp3075.mgmt.esams.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:10 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp7009.magru.wmnet with OS trixie
* 19:09 brett@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp7009.mgmt.magru.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:08 sukhe@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cp3075.esams.wmnet with reason: reimaging
* 19:08 sukhe@cumin1004: START - Cookbook sre.hosts.provision for host cp1103.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:05 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp4046.ulsfo.wmnet with OS trixie
* 19:05 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore1006.eqiad.wmnet with OS bookworm
* 19:04 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3075.*
* 19:02 brett@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp4046.mgmt.ulsfo.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:01 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp1103.eqiad.wmnet
* 19:01 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp1103.eqiad.wmnet
* 19:00 sukhe@cumin1004: START - Cookbook sre.hosts.provision for host cp3075.mgmt.esams.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:58 brett@cumin2003: START - Cookbook sre.hosts.provision for host cp7009.mgmt.magru.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:57 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3075.esams.wmnet
* 18:55 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp1101.eqiad.wmnet
* 18:55 sukhe@puppetserver1001: conftool action : set/weight=1; selector: name=cp1101.eqiad.wmnet
* 18:55 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp4045.ulsfo.wmnet with OS trixie
* 18:52 brett@cumin2003: START - Cookbook sre.hosts.provision for host cp4046.mgmt.ulsfo.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:52 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp1101.eqiad.wmnet with OS trixie
* 18:45 cdobbins@puppetserver1001: conftool action : set/weight=1; selector: name=cp2044.codfw.wmnet
* 18:44 cdobbins@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp2044.codfw.wmnet
* 18:44 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore1006.eqiad.wmnet with reason: host reimage
* 18:40 brett@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp4045.mgmt.ulsfo.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:40 eevans@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore1006.eqiad.wmnet with reason: host reimage
* 18:39 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp3074.esams.wmnet
* 18:36 sukhe@cumin1004: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) config_reloading P<nowiki>{</nowiki>lvs3008.esams.wmnet<nowiki>}</nowiki> and A:liberica
* 18:36 sukhe@cumin1004: START - Cookbook sre.loadbalancer.admin config_reloading P<nowiki>{</nowiki>lvs3008.esams.wmnet<nowiki>}</nowiki> and A:liberica
* 18:33 sukhe@puppetserver1001: conftool action : set/weight=1; selector: name=cp3074.esams.wmnet
* 18:32 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp2044.codfw.wmnet with OS trixie
* 18:31 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp3074.esams.wmnet with OS trixie
* 18:30 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp1101.eqiad.wmnet with reason: host reimage
* 18:29 brett@cumin2003: START - Cookbook sre.hosts.provision for host cp4045.mgmt.ulsfo.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:26 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cp1101.eqiad.wmnet with reason: host reimage
* 18:22 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp7009.magru.wmnet with OS trixie
* 18:20 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1006.eqiad.wmnet with OS bookworm
* 18:20 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:19 eevans@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:19 eevans@cumin1004: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore1006.eqiad.wmnet
* 18:19 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore1006.eqiad.wmnet
* 18:10 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp1101.eqiad.wmnet with OS trixie
* 18:09 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp2044.codfw.wmnet with reason: host reimage
* 18:09 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp4045.ulsfo.wmnet with OS trixie
* 18:08 eevans@cumin1004: START - Cookbook sre.hosts.reboot-single for host sessionstore1006.eqiad.wmnet
* 18:07 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp3074.esams.wmnet with reason: host reimage
* 17:52 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp7009.magru.wmnet with reason: host reimage
* 17:48 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp2044.codfw.wmnet with OS trixie
* 17:43 cdobbins@puppetserver1001: conftool action : set/pooled=no; selector: name=cp2044.codfw.wmnet
* 17:36 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp3074.esams.wmnet with OS trixie
* 17:34 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp4046.ulsfo.wmnet with OS trixie
* 17:34 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp4045.ulsfo.wmnet with OS trixie
* 17:33 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3074.esams.wmnet
* 17:28 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp4046.*
* 17:28 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp4045.*
* 17:26 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp7009.magru.wmnet with OS trixie
* 17:25 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp1101.eqiad.wmnet with OS trixie
* 17:24 cdobbins@puppetserver1001: conftool action : set/pooled=no; selector: name=cp7009.magru.wmnet
* 17:23 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp1101.eqiad.wmnet
* 17:22 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:22 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org
* 17:02 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp7009.*
* 17:01 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:01 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:00 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:59 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:59 vriley@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host ms-be1099
* 16:59 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ms-be1099
* 16:59 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:56 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 16:55 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:55 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:55 vriley@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host ms-be1099
* 16:55 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ms-be1099
* 16:55 vriley@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host ms-be1099
* 16:54 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ms-be1099
* 16:54 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1098.eqiad.wmnet with OS bullseye
* 16:53 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:53 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [ms-be1099] - vriley@cumin1003"
* 16:53 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [ms-be1099] - vriley@cumin1003"
* 16:49 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 16:33 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS bullseye
* 16:21 mutante: temp disabling puppet on C:zookeeper (32 hosts) - safe deploy of https://gerrit.wikimedia.org/r/c/operations/puppet/+/1327569
* 16:04 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341890{{!}}Restore table borders for client-side MathJax (T435274)]], [[gerrit:1340558{{!}}lift IP cap for edit-a-thon /workshop (T437609 T437594 T437470)]] (duration: 24m 19s)
* 15:59 krinkle@deploy1003: anzx, krinkle: Continuing with deployment
* 15:44 krinkle@deploy1003: anzx, krinkle: Backport for [[gerrit:1341890{{!}}Restore table borders for client-side MathJax (T435274)]], [[gerrit:1340558{{!}}lift IP cap for edit-a-thon /workshop (T437609 T437594 T437470)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:40 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1341890{{!}}Restore table borders for client-side MathJax (T435274)]], [[gerrit:1340558{{!}}lift IP cap for edit-a-thon /workshop (T437609 T437594 T437470)]]
* 15:34 elukey@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 15:34 elukey@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 15:33 brennen@deploy1003: Finished deploy [phabricator/deployment@c386249]: deploy phab1005 for [[phab:T437930|T437930]] (duration: 00m 39s)
* 15:33 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS bullseye
* 15:33 brennen@deploy1003: Started deploy [phabricator/deployment@c386249]: deploy phab1005 for [[phab:T437930|T437930]]
* 15:32 elukey@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 15:32 elukey@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'sync'.
* 15:32 brennen@deploy1003: Finished deploy [phabricator/deployment@c386249]: deploy phab2003 for [[phab:T437930|T437930]] (duration: 00m 52s)
* 15:32 elukey@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 15:32 elukey@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'sync'.
* 15:31 brennen@deploy1003: Started deploy [phabricator/deployment@c386249]: deploy phab2003 for [[phab:T437930|T437930]]
* 15:31 elukey@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'.
* 15:31 elukey@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'.
* 15:26 jelto@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on phab2003.codfw.wmnet,phab[1005-1006].eqiad.wmnet with reason: Phabricator deploy
* 15:26 moritzm: pruned obsolete Bullseye image amd-gpu-tester from the docker registry [[phab:T416452|T416452]]
* 15:12 elukey@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 15:12 elukey@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 15:11 elukey@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 15:11 elukey@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 15:00 tgr@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341292{{!}}CommonSettings: Use a restrictive CSP for auth.wikimedia.org (T419684)]] (duration: 25m 11s)
* 14:55 tgr@deploy1003: tgr, arendpieter: Continuing with deployment
* 14:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:53 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:53 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:52 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:49 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 14:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 14:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 14:47 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 14:47 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 14:45 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 14:45 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 14:45 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 14:44 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 14:42 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 14:42 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 14:39 tgr@deploy1003: tgr, arendpieter: Backport for [[gerrit:1341292{{!}}CommonSettings: Use a restrictive CSP for auth.wikimedia.org (T419684)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:34 tgr@deploy1003: Started scap sync-world: Backport for [[gerrit:1341292{{!}}CommonSettings: Use a restrictive CSP for auth.wikimedia.org (T419684)]]
* 14:17 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341895{{!}}ReportIncidentController: Instance cache expensive methods (T437588)]] (duration: 11m 56s)
* 14:16 btullis@cumin1004: END (PASS) - Cookbook sre.ceph.rotate-osd-keys (exit_code=0) rolling rotate_keys on A:cephosd-codfw
* 14:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:12 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 14:09 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1341895{{!}}ReportIncidentController: Instance cache expensive methods (T437588)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:05 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1341895{{!}}ReportIncidentController: Instance cache expensive methods (T437588)]]
* 13:43 btullis@cumin1004: START - Cookbook sre.ceph.rotate-osd-keys rolling rotate_keys on A:cephosd-codfw
* 13:36 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341861{{!}}SuggestedInvestigations: Update "sockpuppet" queue view defaults (T438018)]] (duration: 10m 23s)
* 13:32 stran@deploy1003: stran: Continuing with deployment
* 13:30 stran@deploy1003: stran: Backport for [[gerrit:1341861{{!}}SuggestedInvestigations: Update "sockpuppet" queue view defaults (T438018)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:26 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1341861{{!}}SuggestedInvestigations: Update "sockpuppet" queue view defaults (T438018)]]
* 13:21 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:20 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:19 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334946{{!}}ArticleGuidance: Remove the experiment configuration keys (T434487)]] (duration: 09m 19s)
* 13:16 btullis@cumin1004: END (PASS) - Cookbook sre.ceph.rotate-osd-keys (exit_code=0) rolling rotate_keys on P<nowiki>{</nowiki>cephosd2001.codfw.wmnet<nowiki>}</nowiki> and (A:cephosd-codfw or A:cephosd-eqiad)
* 13:15 sbisson@deploy1003: sbisson: Continuing with deployment
* 13:14 sbisson@deploy1003: sbisson: Backport for [[gerrit:1334946{{!}}ArticleGuidance: Remove the experiment configuration keys (T434487)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) config_reloading P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:10 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1334946{{!}}ArticleGuidance: Remove the experiment configuration keys (T434487)]]
* 13:10 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin config_reloading P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:08 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:08 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:07 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:07 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:07 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:07 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:03 btullis@cumin1004: START - Cookbook sre.ceph.rotate-osd-keys rolling rotate_keys on P<nowiki>{</nowiki>cephosd2001.codfw.wmnet<nowiki>}</nowiki> and (A:cephosd-codfw or A:cephosd-eqiad)
* 13:00 btullis@cumin1004: END (PASS) - Cookbook sre.ceph.rotate-osd-keys (exit_code=0) rolling rotate_keys on P<nowiki>{</nowiki>cephosd2001.codfw.wmnet<nowiki>}</nowiki> and (A:cephosd-codfw or A:cephosd-eqiad)
* 12:59 btullis@cumin1004: START - Cookbook sre.ceph.rotate-osd-keys rolling rotate_keys on P<nowiki>{</nowiki>cephosd2001.codfw.wmnet<nowiki>}</nowiki> and (A:cephosd-codfw or A:cephosd-eqiad)
* 12:46 btullis@cumin1004: END (PASS) - Cookbook sre.ceph.rotate-osd-keys (exit_code=0) rolling rotate_keys on P<nowiki>{</nowiki>cephosd2001.codfw.wmnet<nowiki>}</nowiki> and (A:cephosd-codfw or A:cephosd-eqiad)
* 12:45 btullis@cumin1004: START - Cookbook sre.ceph.rotate-osd-keys rolling rotate_keys on P<nowiki>{</nowiki>cephosd2001.codfw.wmnet<nowiki>}</nowiki> and (A:cephosd-codfw or A:cephosd-eqiad)
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: network maintenance finished, [[phab:T437984|T437984]]]
* 12:40 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: network maintenance finished, [[phab:T437984|T437984]]]
* 12:29 XioNoX: asw1-b4-magru> request system reboot - [[phab:T437984|T437984]]
* 12:24 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 12:24 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 12:24 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 12:23 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 12:23 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 12:23 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 12:22 moritzm: installing shadow security updates
* 12:19 slyngshede@puppetserver1001: conftool action : set/weight=1; selector: name=cp7010.magru.wmnet
* 12:13 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 12:13 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 12 hosts with reason: Switch maintenance
* 12:12 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-b4-magru,asw1-b4-magru IPv6,asw1-b4-magru.mgmt with reason: Switch maintenance
* 12:11 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 12:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: switch reboot, [[phab:T437984|T437984]]]
* 12:11 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: switch reboot, [[phab:T437984|T437984]]]
* 12:09 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on install7002.wikimedia.org with reason: switch reboot
* 12:08 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 12:07 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 12:07 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 12:07 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 12:06 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 12:06 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 12:03 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 12:03 XioNoX: push pfw policies - [[phab:T437627|T437627]]
* 12:01 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 12:01 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) config_reloading P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 12:00 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin config_reloading P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 11:56 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 11:56 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 11:33 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2250.codfw.wmnet with reason: cloning db2201
* 11:18 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7004.magru.wmnet
* 11:17 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7004.magru.wmnet
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 11:05 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 11:05 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 11:04 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 10:51 slyngshede@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp7010.magru.wmnet
* 10:34 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 10:33 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 10:30 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1098.eqiad.wmnet with OS trixie
* 10:21 moritzm: failover Ganeti master in magru to ganeti7001
* 10:20 moritzm: increased DRBD replication speed in Ganeti/magru [[phab:T428878|T428878]]
* 10:10 hashar@deploy1003: Finished deploy [integration/docroot@5cf09c8]: build: Updating npm dependencies (duration: 00m 13s)
* 10:10 hashar@deploy1003: Started deploy [integration/docroot@5cf09c8]: build: Updating npm dependencies
* 10:09 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 10:08 moritzm: increased DRBD replication speed in Ganeti/esams [[phab:T428878|T428878]]
* 10:07 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 10:05 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 10:05 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:39 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: switches reboot, [[phab:T437984|T437984]]]
* 09:39 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool esams [reason: switches reboot, [[phab:T437984|T437984]]]
* 09:31 XioNoX: asw1-by27-esams> request system reboot - [[phab:T437984|T437984]]
* 09:30 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS trixie
* 09:28 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp7010.magru.wmnet with OS trixie
* 09:26 ayounsi@cumin1003: END (FAIL) - Cookbook sre.network.depool-rack (exit_code=99) with action 'depool' for esams rack BY27
* 09:24 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for esams rack BY27
* 09:24 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1098.eqiad.wmnet with OS trixie
* 09:23 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS trixie
* 09:22 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be1098.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 09:15 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.20 refs [[phab:T430839|T430839]]
* 09:10 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1098.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 09:06 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-by27-esams,asw1-by27-esams IPv6,asw1-by27-esams.mgmt with reason: Switch maintenance
* 09:05 ayounsi@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on asw1-by27-esams IPv6,asw1-by27-esams.mgmt,asw1-by-27-esams with reason: Switch maintenance
* 09:04 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 12 hosts with reason: Switch maintenance
* 09:04 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp7010.magru.wmnet with reason: host reimage
* 09:01 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: switches reboot, [[phab:T437984|T437984]]]
* 09:00 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: switches reboot, [[phab:T437984|T437984]]]
* 09:00 slyngshede@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp7010.magru.wmnet with reason: host reimage
* 08:59 jnuche@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341698{{!}}RestSandbox: Pass JsonLocalizer instead of ResponseFactory to ModuleManager (T437982)]] (duration: 12m 03s)
* 08:55 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2197.codfw.wmnet with reason: cloning db2201
* 08:55 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on install3004.wikimedia.org with reason: switch reboot
* 08:53 jnuche@deploy1003: jnuche: Continuing with deployment
* 08:52 jnuche@deploy1003: jnuche: Backport for [[gerrit:1341698{{!}}RestSandbox: Pass JsonLocalizer instead of ResponseFactory to ModuleManager (T437982)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:49 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: sync
* 08:49 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: sync
* 08:47 jnuche@deploy1003: Started scap sync-world: Backport for [[gerrit:1341698{{!}}RestSandbox: Pass JsonLocalizer instead of ResponseFactory to ModuleManager (T437982)]]
* 08:33 slyngshede@cumin1003: START - Cookbook sre.hosts.reimage for host cp7010.magru.wmnet with OS trixie
* 08:26 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be1098.eqiad.wmnet with OS trixie
* 08:26 slyngshede@puppetserver1001: conftool action : set/pooled=no; selector: name=cp7010.magru.wmnet
* 08:25 XioNoX: asw1-b3-magru - Disable logging and file logging for BRCM_PKT - [[phab:T437984|T437984]]
* 08:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS trixie
* 08:18 dpogorzelski@dns1004: END - running authdns-update
* 08:15 dpogorzelski@dns1004: START - running authdns-update
* 08:14 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3005.esams.wmnet
* 08:13 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3005.esams.wmnet
* 08:07 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341274{{!}}SI: Implement "queue view" functionality (T437183)]], [[gerrit:1341242{{!}}SuggestedInvestigations: Add and enable 'sockpuppets' queue view (T437183)]], [[gerrit:1341254{{!}}Add wmf-specific Special:SuggestedInvestigations messages (T437183)]] (duration: 55m 27s)
* 07:54 stran@deploy1003: stran: Continuing with deployment
* 07:31 stran@deploy1003: stran: Backport for [[gerrit:1341274{{!}}SI: Implement "queue view" functionality (T437183)]], [[gerrit:1341242{{!}}SuggestedInvestigations: Add and enable 'sockpuppets' queue view (T437183)]], [[gerrit:1341254{{!}}Add wmf-specific Special:SuggestedInvestigations messages (T437183)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:18 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply
* 07:18 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply
* 07:11 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1341274{{!}}SI: Implement "queue view" functionality (T437183)]], [[gerrit:1341242{{!}}SuggestedInvestigations: Add and enable 'sockpuppets' queue view (T437183)]], [[gerrit:1341254{{!}}Add wmf-specific Special:SuggestedInvestigations messages (T437183)]]
* 07:06 moritzm: pruned obsolete Bullseye image python3-devel from the docker registry [[phab:T416452|T416452]]
* 06:51 moritzm: pruned obsolete Bullseye image python3-build-bullseye from the docker registry [[phab:T416452|T416452]]
* 05:59 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply
* 05:49 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply
* 05:48 oblivian@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 05:47 oblivian@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 05:38 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply
* 05:28 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-video: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-video: apply
* 05:10 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-video: apply
* 05:10 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-video: apply
* 05:10 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-video: apply
* 05:10 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-video: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-syntaxhighlight: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-syntaxhighlight: apply
* 05:10 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-syntaxhighlight: apply
* 05:10 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-syntaxhighlight: apply
* 05:10 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-syntaxhighlight: apply
* 05:10 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-syntaxhighlight: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-media: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-media: apply
* 05:10 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-media: apply
* 05:10 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-media: apply
* 05:10 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-media: apply
* 05:10 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-media: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-constraints: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-constraints: apply
* 05:10 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-constraints: apply
* 05:09 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-constraints: apply
* 05:09 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-constraints: apply
* 05:09 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-constraints: apply
* 05:08 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply
* 05:08 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox: apply
* 05:07 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply
* 05:07 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox: apply
* 05:07 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox: apply
* 05:07 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox: apply
* 05:07 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply
* 05:07 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply
* 05:06 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-timeline: apply
* 05:06 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-timeline: apply
* 05:06 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply
* 05:06 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply
* 04:07 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.17 (duration: 07m 10s)
* 03:06 mwpresync@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.19,1.47.0-wmf.20,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.20 refs [[phab:T430839|T430839]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 22s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:43 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 00:23 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore1005.eqiad.wmnet with reason: host reimage
* 00:19 jclark@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore1005.eqiad.wmnet with reason: host reimage
* 00:17 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 00:07 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sessionstore1005.eqiad.wmnet with OS bookworm
== 2026-09-14 ==
* 23:41 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 23:26 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324966{{!}}Enable Produnto on pilot wikis (T421436)]] (duration: 12m 59s)
* 23:22 tstarling@deploy1003: tstarling: Continuing with deployment
* 23:17 tstarling@deploy1003: tstarling: Backport for [[gerrit:1324966{{!}}Enable Produnto on pilot wikis (T421436)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:13 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1324966{{!}}Enable Produnto on pilot wikis (T421436)]]
* 23:01 eevans@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 22:41 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341385{{!}}VisualEditor: don't register settings tool in wikitextCommandRegistry (T437810)]] (duration: 09m 22s)
* 22:36 kemayo@deploy1003: kemayo: Continuing with deployment
* 22:36 kemayo@deploy1003: kemayo: Backport for [[gerrit:1341385{{!}}VisualEditor: don't register settings tool in wikitextCommandRegistry (T437810)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:31 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1341385{{!}}VisualEditor: don't register settings tool in wikitextCommandRegistry (T437810)]]
* 22:22 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 22:22 eevans@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 21:43 sbassett: Deployed security fix for [[phab:T435623|T435623]]
* 21:29 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1098.eqiad.wmnet with OS bullseye
* 21:29 sbassett: Deployed security fix for [[phab:T434372|T434372]]
* 21:26 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 21:26 eevans@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 21:05 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338999{{!}}Enable ReadingLists for all logged-in users on English Wikipedia (T434923)]], [[gerrit:1340004{{!}}Enable Reading Recommendations experiment on test wiki (T437665)]] (duration: 11m 03s)
* 21:00 aude@deploy1003: aude, jdlrobson: Continuing with deployment
* 20:58 aude@deploy1003: aude, jdlrobson: Backport for [[gerrit:1338999{{!}}Enable ReadingLists for all logged-in users on English Wikipedia (T434923)]], [[gerrit:1340004{{!}}Enable Reading Recommendations experiment on test wiki (T437665)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:54 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1338999{{!}}Enable ReadingLists for all logged-in users on English Wikipedia (T434923)]], [[gerrit:1340004{{!}}Enable Reading Recommendations experiment on test wiki (T437665)]]
* 20:47 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341278{{!}}[A11y] Add list semantics to ReadingList page (T435864 T434923)]] (duration: 12m 49s)
* 20:43 aude@deploy1003: aude, jdlrobson: Continuing with deployment
* 20:39 aude@deploy1003: aude, jdlrobson: Backport for [[gerrit:1341278{{!}}[A11y] Add list semantics to ReadingList page (T435864 T434923)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:34 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1341278{{!}}[A11y] Add list semantics to ReadingList page (T435864 T434923)]]
* 20:32 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339811{{!}}wmf-config: Register content/v2-beta REST module as disabled (T432798)]], [[gerrit:1338274{{!}}wikifunctions: Move abstract fragments to mainstash (T432849)]] (duration: 25m 21s)
* 20:28 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:27 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:27 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:27 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:25 jforrester@deploy1003: jforrester, aghirelli: Continuing with deployment
* 20:24 jforrester@deploy1003: jforrester, aghirelli: Backport for [[gerrit:1339811{{!}}wmf-config: Register content/v2-beta REST module as disabled (T432798)]], [[gerrit:1338274{{!}}wikifunctions: Move abstract fragments to mainstash (T432849)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS bullseye
* 20:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be1098.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 20:06 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1339811{{!}}wmf-config: Register content/v2-beta REST module as disabled (T432798)]], [[gerrit:1338274{{!}}wikifunctions: Move abstract fragments to mainstash (T432849)]]
* 20:04 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1098.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 20:03 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be1098
* 20:02 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ms-be1098
* 20:02 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 20:02 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [ms-be1098] - vriley@cumin1003"
* 20:02 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [ms-be1098] - vriley@cumin1003"
* 19:59 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 19:58 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 19:57 dzahn@dns1005: END - running authdns-update
* 19:55 eevans@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 19:55 dzahn@dns1005: START - running authdns-update
* 19:54 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 19:38 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341273{{!}}[CodeMirror] enable for new users (enwiki), new VE integration (global) (T288161 T432558)]] (duration: 33m 51s)
* 19:26 musikanimal@deploy1003: musikanimal: Continuing with deployment
* 19:22 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1341273{{!}}[CodeMirror] enable for new users (enwiki), new VE integration (global) (T288161 T432558)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:04 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1341273{{!}}[CodeMirror] enable for new users (enwiki), new VE integration (global) (T288161 T432558)]]
* 18:53 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:50 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 18:50 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 18:49 eevans@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 18:26 brett@cumin2003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d6-eqiad
* 18:26 brett@cumin2003: START - Cookbook sre.network.tls for network device lsw1-d6-eqiad
* 18:26 brett@cumin2003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-eqiad
* 18:26 brett@cumin2003: START - Cookbook sre.network.tls for network device ssw1-d8-eqiad
* 18:25 brett@cumin2003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-eqiad
* 18:25 brett@cumin2003: START - Cookbook sre.network.tls for network device lsw1-d4-eqiad
* 18:25 root@cumin2003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-eqiad
* 18:25 root@cumin2003: START - Cookbook sre.network.tls for network device lsw1-d2-eqiad
* 18:17 jclark@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:14 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337611{{!}}extension-list: Add ModeratorToolkit (T431000)]] (duration: 09m 34s)
* 18:10 samtar@deploy1003: samtar: Continuing with deployment
* 18:09 samtar@deploy1003: samtar: Backport for [[gerrit:1337611{{!}}extension-list: Add ModeratorToolkit (T431000)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:06 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 18:05 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1337611{{!}}extension-list: Add ModeratorToolkit (T431000)]]
* 18:01 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 17:48 jclark@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 17:07 tgr@deploy1003: Finished scap sync-world: Backport for [[gerrit:1330446{{!}}CommonSettings: Use a restrictive, eval-free CSP for auth.wikimedia.org (T419684)]] (duration: 23m 19s)
* 17:00 tgr@deploy1003: arendpieter, tgr: Rolling back deployment
* 16:52 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 16:51 eevans@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 16:49 tgr@deploy1003: arendpieter, tgr: Backport for [[gerrit:1330446{{!}}CommonSettings: Use a restrictive, eval-free CSP for auth.wikimedia.org (T419684)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:44 tgr@deploy1003: Started scap sync-world: Backport for [[gerrit:1330446{{!}}CommonSettings: Use a restrictive, eval-free CSP for auth.wikimedia.org (T419684)]]
* 16:09 Amir1: drop links tables from db1252 ([[phab:T437278|T437278]])
* 16:07 Amir1: drop links tables from db2240 ([[phab:T437278|T437278]])
* 16:05 Amir1: drop non-links tables from db2247 ([[phab:T437278|T437278]])
* 15:53 Lucas_WMDE: UTC afternoon backport+config window belatedly done
* 15:50 lucaswerkmeister-wmde@deploy1003: mwscript-k8s job started: namespaceDupes abstractwiki --fix # [[phab:T437772|T437772]]
* 15:49 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335401{{!}}Adjust extendedconfirmed calculation to first edit on viwiki (T437006)]], [[gerrit:1340505{{!}}core-Namespaces: Add AW and AWT alias for its talk in abstractwiki (T437772)]] (duration: 10m 23s)
* 15:48 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 15:47 eevans@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 15:45 lucaswerkmeister-wmde@deploy1003: bunnypranav, lucaswerkmeister-wmde, tryvix1509: Continuing with deployment
* 15:43 lucaswerkmeister-wmde@deploy1003: bunnypranav, lucaswerkmeister-wmde, tryvix1509: Backport for [[gerrit:1335401{{!}}Adjust extendedconfirmed calculation to first edit on viwiki (T437006)]], [[gerrit:1340505{{!}}core-Namespaces: Add AW and AWT alias for its talk in abstractwiki (T437772)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:39 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1335401{{!}}Adjust extendedconfirmed calculation to first edit on viwiki (T437006)]], [[gerrit:1340505{{!}}core-Namespaces: Add AW and AWT alias for its talk in abstractwiki (T437772)]]
* 15:36 elukey@dns1004: END - running authdns-update
* 15:36 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 15:35 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 15:35 moritzm: installing shadow security updates
* 15:35 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 15:34 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:33 elukey@dns1004: START - running authdns-update
* 15:33 eevans@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:32 eevans@cumin1004: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore1005.eqiad.wmnet
* 15:29 lucaswerkmeister-wmde@deploy1003: mwscript-k8s job started: namespaceDupes afwiki --fix # [[phab:T437576|T437576]]
* 15:29 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338902{{!}}afwiki: Create Draft and Draft talk namespaces (T437576)]] (duration: 15m 44s)
* 15:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 15:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 15:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 15:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 15:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 15:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 15:21 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, tryvix1509: Continuing with deployment
* 15:21 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:17 eevans@cumin1004: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore1005.eqiad.wmnet
* 15:17 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, tryvix1509: Backport for [[gerrit:1338902{{!}}afwiki: Create Draft and Draft talk namespaces (T437576)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:17 eevans@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore1005.eqiad.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 15:13 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore1004.eqiad.wmnet
* 15:13 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1338902{{!}}afwiki: Create Draft and Draft talk namespaces (T437576)]]
* 15:06 eevans@cumin1004: START - Cookbook sre.hosts.reboot-single for host sessionstore1004.eqiad.wmnet
* 14:59 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore1004.eqiad.wmnet with OS bookworm
* 14:38 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore1004.eqiad.wmnet with reason: host reimage
* 14:33 marostegui@dns1004: END - running authdns-update
* 14:32 eevans@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore1004.eqiad.wmnet with reason: host reimage
* 14:30 marostegui@dns1004: START - running authdns-update
* 14:15 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1004.eqiad.wmnet with OS bookworm
* 14:14 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore1004.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:13 eevans@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1004.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:13 eevans@cumin1004: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore1004.eqiad.wmnet
* 14:13 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore1004.eqiad.wmnet
* 14:05 moritzm: kick off a rebuild of base images on build2004
* 14:05 moritzm: kick off a rebuild of base images on build2004
* 14:00 eevans@cumin1004: START - Cookbook sre.hosts.reboot-single for host sessionstore1004.eqiad.wmnet
* 14:00 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:00 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 13:56 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:54 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:52 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:44 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 13:43 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 13:42 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:40 eevans@cumin1004: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore1004.eqiad.wmnet
* 13:40 eevans@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore1004.eqiad.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 13:40 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:40 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:39 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:39 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:39 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 13:38 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 13:38 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:38 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:38 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 13:37 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 13:37 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:36 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:36 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 13:35 oblivian@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:35 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 13:35 oblivian@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:35 oblivian@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:34 oblivian@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'.
* 13:34 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 13:34 oblivian@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:34 oblivian@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:34 oblivian@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:34 oblivian@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 13:34 oblivian@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'.
* 13:33 oblivian@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'.
* 13:33 oblivian@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
* 13:33 oblivian@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
* 13:33 oblivian@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 13:33 oblivian@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 13:33 oblivian@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 13:32 oblivian@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 13:32 oblivian@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 13:32 oblivian@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 13:32 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 13:32 oblivian@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 13:32 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 13:32 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 13:29 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:23 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1003.eqiad.wmnet
* 13:21 sukhe: sudo cumin -b11 "A:cp-text" "run-puppet-agent --enable 'merging CR 1338134'" [[phab:T425441|T425441]]
* 13:20 sukhe: sudo cumin -b11 "A:cp-text" "run-puppet-agent --enable 'merging CR 1338134'"[[phab:T425441|T425441]]
* 13:19 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:17 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1003.eqiad.wmnet
* 13:14 moritzm: installing Bird security updates
* 13:09 sukhe: sudo cumin "A:cp-text" "disable-puppet 'merging CR 1338134'"
* 13:06 jmm@dns1004: END - running authdns-update
* 13:04 jmm@dns1004: START - running authdns-update
* 12:58 moritzm: update Trixie installer image to 13.7 [[phab:T437715|T437715]]
* 12:58 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 12:54 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 12:52 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 12:49 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 12:48 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 12:47 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 12:44 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 12:44 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 12:42 oblivian@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 12:40 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 12:40 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 12:40 oblivian@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 12:39 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 12:39 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 12:39 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 12:37 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 12:37 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 12:36 marostegui@cumin1004: conftool action : set/pooled=yes; selector: name=clouddb1025.eqiad.wmnet,service=x4
* 12:34 _joe_: adding gvisor labels to all wikikube clusters nodes
* 12:30 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 12:14 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 12:14 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 11:22 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1260: After cloning
* 10:48 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:45 ladsgroup@dns1004: END - running authdns-update
* 10:42 ladsgroup@dns1004: START - running authdns-update
* 10:37 marostegui@cumin1004: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=x4
* 10:37 marostegui@cumin1004: START - Cookbook sre.mysql.pool pool db1260: After cloning
* 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 10:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 10:04 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 10:04 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 09:53 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:53 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:41 Amir1: drop links tables from db2172 ([[phab:T437278|T437278]])
* 09:40 Amir1: drop links tables from db1228 ([[phab:T437278|T437278]])
* 09:08 marostegui: Stop mariadb on db1260 to clone dbstore1007, there will be lag on wikireplicas:x4 https://phabricator.wikimedia.org/T437839
* 09:07 taavi@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1025.eqiad.wmnet
* 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1260: Needs to clone another host from this one
* 09:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1260: Needs to clone another host from this one
* 09:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb[1024-1025].eqiad.wmnet,db[1155,1260].eqiad.wmnet,dbstore1007.eqiad.wmnet with reason: Adding x4
* 08:44 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:42 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:41 moritzm: pruned obsolete Bullseye images php8.3-icu72-cli / php8.3-icu72-fpm-multiversion-base / php8.3-icu72-fpm from the docker registry [[phab:T416452|T416452]]
* 08:37 taavi@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet
* 08:37 taavi@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1024.eqiad.wmnet
* 08:36 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on dbstore1007.eqiad.wmnet with reason: Adding x4
* 08:35 moritzm: pruned obsolete Bullseye images php8.1-cli/php8.1-fpm/ php8.1-fpm-multiversion-base from the docker registry [[phab:T416452|T416452]]
* 08:10 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: sync
* 08:08 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: sync
* 07:58 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 23s)
* 07:57 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 07:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 07:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 07:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 07:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 07:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 07:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 07:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 07:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 07:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 07:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 07:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 07:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 07:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 07:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 07:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 07:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 07:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 07:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 07:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 07:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 07:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 07:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 07:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 07:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 07:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 07:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 07:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 07:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 07:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 07:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 07:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 07:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 07:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 07:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 07:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 07:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 07:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 07:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 07:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 07:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 07:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 07:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 07:31 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1340801{{!}}Allow title-like strings with Package: prefix in require() (T430644)]], [[gerrit:1340802{{!}}Runtime: Add a facility for loading files by title (T430644)]] (duration: 34m 30s)
* 07:18 tstarling@deploy1003: tstarling: Continuing with deployment
* 07:17 tstarling@deploy1003: tstarling: Backport for [[gerrit:1340801{{!}}Allow title-like strings with Package: prefix in require() (T430644)]], [[gerrit:1340802{{!}}Runtime: Add a facility for loading files by title (T430644)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 06:56 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1340801{{!}}Allow title-like strings with Package: prefix in require() (T430644)]], [[gerrit:1340802{{!}}Runtime: Add a facility for loading files by title (T430644)]]
* 06:26 TimStarling: on deploy1003: docker image pull docker-registry.wikimedia.org/php8.3-fpm-multiversion-base
* 05:51 _joe_: pulled bookworm:latest from build2004 to build2001 [[phab:T437829|T437829]]
* 05:39 _joe_: force-running build-base-images on build2004 for [[phab:T437829|T437829]]
* 04:53 TimStarling: on build2001 rebuilding base images [[phab:T437829|T437829]]
* 03:00 tstarling@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.18,1.47.0-wmf.19,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted
* 02:59 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1340801{{!}}Allow title-like strings with Package: prefix in require() (T430644)]], [[gerrit:1340802{{!}}Runtime: Add a facility for loading files by title (T430644)]]
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-13 ==
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 29s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-12 ==
* 19:40 ladsgroup@cumin1003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-eqiad
* 19:32 ladsgroup@cumin1003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-eqiad
* 19:30 ladsgroup@cumin1003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw
* 19:21 ladsgroup@cumin1003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 35s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-11 ==
* 21:51 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 21:50 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 16:47 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339808{{!}}Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813{{!}}Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]] (duration: 07m 23s)
* 16:43 sbassett@deploy1003: sbassett: Continuing with deployment
* 16:42 sbassett@deploy1003: sbassett: Backport for [[gerrit:1339808{{!}}Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813{{!}}Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:40 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1339808{{!}}Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813{{!}}Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]]
* 16:08 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 14:39 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 14:10 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:10 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:10 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:09 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 13:40 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:28 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:11 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:10 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: db1199 repool
* 11:05 moritzm: installing Linux 6.1.187 on Bookworm hosts
* 11:05 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Jcrespo out of all services on: 2443 hosts
* 10:44 aokoth@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2020 in turn
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2019 in turn
* 10:41 Emperor: restart versitygw@objectstorage0[0-3].service on backup2018 in turn
* 10:40 Emperor: restart versitygw@objectstorage0[0-3].service on backup2017 in turn
* 10:39 Emperor: restart versitygw@objectstorage0[0-3].service on backup2016 in turn
* 10:37 Emperor: restart versitygw@objectstorage0[0-3].service on backup2015 in turn
* 10:36 Emperor: restart versitygw@objectstorage0[0-3].service on backup1020 in turn
* 10:35 Emperor: restart versitygw@objectstorage0[0-3].service on backup1019 in turn
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1199: db1199 repool
* 10:33 Emperor: restart versitygw@objectstorage0[0-3].service on backup1018 in turn
* 10:32 Emperor: restart versitygw@objectstorage0[0-3].service on backup1017 in turn
* 10:30 Emperor: restart versitygw@objectstorage0[0-3].service on backup1016 in turn
* 10:20 Emperor: restart versitygw@objectstorage0[1-3].service on backup1015 in turn
* 10:17 Emperor: restart versitygw@objectstorage00.service on backup1015
* 10:15 aokoth@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 08:46 slyngs: Update CAS/SSO to CAS 7.3.8.3
* 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:45 slyngshede@dns1004: END - running authdns-update
* 08:43 slyngshede@dns1004: START - running authdns-update
* 08:36 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:28 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 7 hosts with reason: Restarting s5
* 08:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db[1154,1269].eqiad.wmnet with reason: Restarting s5
* 08:20 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:20 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Repooling db1159
* 07:58 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:58 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:54 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:28 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:27 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1199: Needs to clone another host from this one
* 07:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1199: Needs to clone another host from this one
* 07:16 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Repooling db1159
* 07:15 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1199.eqiad.wmnet with reason: Cloning s4
* 07:10 TimStarling: killed jobs for [[phab:T437056|T437056]] since they weren't purging
* 06:38 TimStarling: also started refreshLinks for ptwiki and zhwiki, reparsing ~3000 pages altogether [[phab:T437056|T437056]]
* 06:27 TimStarling: for [[phab:T437056|T437056]]: mwscript-k8s refreshLinks.php --wiki=eswiki --tracking-category scribunto-common-error-category
* 05:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1159: Needs to clone another host from this one
* 05:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1159: Needs to clone another host from this one
* 05:30 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1159.eqiad.wmnet with reason: Cloning
* 05:29 marostegui: Start cloning db1245:s5 [[phab:T437563|T437563]]
* 05:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet,db1245.eqiad.wmnet with reason: Cloning
* 05:25 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] (duration: 09m 59s)
* 05:21 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:20 tstarling@deploy1003: tstarling: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:15 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 50s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-10 ==
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1349.eqiad.wmnet
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1349.eqiad.wmnet
* 23:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1349.eqiad.wmnet
* 23:09 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1349
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1349
* 22:32 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1349
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:28 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1349
* 22:27 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:27 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1349.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1348.eqiad.wmnet
* 22:23 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 22:12 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] (duration: 10m 59s)
* 22:06 musikanimal@deploy1003: kemayo, musikanimal: Rolling back deployment
* 22:05 musikanimal@deploy1003: kemayo, musikanimal: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:01 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]]
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:00 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:52 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:47 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:47 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] (duration: 13m 23s)
* 21:46 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:46 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2006.codfw.wmnet
* 21:42 derenrich@deploy1003: derenrich: Continuing with deployment
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1348
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1348
* 21:39 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1348
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:39 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:37 derenrich@deploy1003: derenrich: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:35 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:34 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1348
* 21:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 21:33 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1348.eqiad.wmnet
* 21:33 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]]
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1348.eqiad.wmnet
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1348.eqiad.wmnet
* 21:31 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2006.codfw.wmnet
* 21:31 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] (duration: 09m 45s)
* 21:27 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:26 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]]
* 21:19 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:19 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2006.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 21:17 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] (duration: 13m 54s)
* 21:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2005.codfw.wmnet with OS bookworm
* 21:12 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:07 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdeb
* 21:03 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]]
* 20:54 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:52 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] (duration: 23m 49s)
* 20:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:47 jdrewniak@deploy1003: jdrewniak, milazg: Continuing with deployment
* 20:34 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1346.eqiad.wmnet
* 20:32 jdrewniak@deploy1003: jdrewniak, milazg: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2005.codfw.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:28 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]]
* 20:27 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2005.codfw.wmnet
* 20:26 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:24 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:22 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] (duration: 11m 24s)
* 20:17 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 20:15 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2005.codfw.wmnet
* 20:14 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:12 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 20:11 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:07 jdrewniak@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.18,1.47.0-wmf.19,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted
* 20:05 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:02 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:01 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2005.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 19:53 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:50 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1346
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 19:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1346.eqiad.wmnet
* 19:21 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 thcipriani: Gerrit downtime incoming for upgrade
* 19:17 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hadoop.roll-restart-workers (exit_code=0) restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 dzahn@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:30:00 on gerrit.wikimedia.org with reason: maintenance upgrade
* 19:16 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on gerrit2003.wikimedia.org with reason: maintenance upgrade
* 19:04 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 18:21 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e] (duration: 00m 59s)
* 18:20 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e]
* 18:19 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e] (duration: 05m 13s)
* 18:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 18:14 otto@deploy1003: Started deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e]
* 18:13 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e] (duration: 00m 39s)
* 18:13 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e]
* 18:13 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 18:12 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 18:11 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 18:11 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit2002.wikimedia.org with reason: maintenance upgrade
* 18:11 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:11 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 18:10 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit1003.wikimedia.org with reason: maintenance upgrade
* 18:09 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 18:08 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 18:06 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 16:40 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 16:35 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 16:33 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]] synced to the te
* 16:28 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 15:33 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4
* 15:04 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] (duration: 08m 08s)
* 15:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host clouddumps1001.wikimedia.org with OS bookworm
* 14:59 samtar@deploy1003: samtar: Continuing with deployment
* 14:58 samtar@deploy1003: samtar: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:56 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]]
* 14:40 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:07 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 13:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host clouddumps1001.wikimedia.org with OS bookworm
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:55 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:51 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:42 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:41 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:48 klausman@dns1004: END - running authdns-update
* 12:46 klausman@dns1004: START - running authdns-update
* 12:35 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 12:35 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 12:05 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning x4
* 12:01 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2005.codfw.wmnet
* 11:55 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2005.codfw.wmnet
* 11:54 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1144.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:52 cgoubert@dns1004: END - running authdns-update
* 11:49 cgoubert@dns1004: START - running authdns-update
* 11:31 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2004.codfw.wmnet
* 11:25 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2004.codfw.wmnet
* 11:24 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:16 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1200.eqiad.wmnet
* 11:16 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1200.eqiad.wmnet
* 11:04 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1200.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:04 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1199.eqiad.wmnet
* 11:03 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1199.eqiad.wmnet
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:42 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1199.eqiad.wmnet with reason: Upgrading RAID firmware
* 10:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1024.eqiad.wmnet with reason: Cloning x4
* 10:00 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 09:56 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 09:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1024.eqiad.wmnet
* 09:45 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 09:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 09:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:04 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:51 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:46 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning x4
* 08:43 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:34 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] (duration: 09m 56s)
* 08:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:30 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:29 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:28 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:24 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]]
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 08:09 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 07:58 XioNoX: netflow1004:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:56 XioNoX: netflow2005:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 07:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 07:03 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:59 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:43 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:42 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:39 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 06:38 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 06:37 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 05:04 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 05:03 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tool-server' for release 'main' .
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 38s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1345.eqiad.wmnet
* 00:11 dzahn@dns1004: END - running authdns-update
* 00:08 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 00:08 dzahn@dns1004: START - running authdns-update
== 2026-09-09 ==
* 23:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:29 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1345
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 23:28 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1345.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1344.eqiad.wmnet
* 23:22 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] (duration: 11m 15s)
* 23:18 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 23:16 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 23:11 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]]
* 22:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:51 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2004.codfw.wmnet with OS bookworm
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1344
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1344
* 22:37 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1344
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1344
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 22:32 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] (duration: 10m 21s)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1344.eqiad.wmnet
* 22:27 derenrich@deploy1003: derenrich, egardner: Continuing with deployment
* 22:26 derenrich@deploy1003: derenrich, egardner: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:24 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:22 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]]
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:19 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] (duration: 13m 40s)
* 22:16 derenrich@deploy1003: derenrich: Rolling back deployment
* 22:10 derenrich@deploy1003: derenrich: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:05 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]]
* 22:03 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:44 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:40 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:36 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1343
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1343
* 21:27 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1343
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:27 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:23 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:22 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1343
* 21:22 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] (duration: 12m 29s)
* 21:21 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 21:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1343.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1342.eqiad.wmnet
* 21:17 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:14 jforrester@deploy1003: jforrester: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:09 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]]
* 21:08 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:07 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 21:07 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:06 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2004.codfw.wmnet
* 20:59 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:57 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:56 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2004.codfw.wmnet
* 20:53 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:53 bking@cumin2003: END (ERROR) - Cookbook sre.presto.roll-restart-workers (exit_code=97) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:53 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:50 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* {{safesubst:SAL entry|1=20:45 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2}}
* 20:42 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* 20:41 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 20:40 sbassett@deploy1003: aranyap, sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:39 sbassett@deploy1003: aranyap, sbassett: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "Filter}}
* 20:35 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* {{safesubst:SAL entry|1=20:34 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "}}
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:28 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:24 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 20:24 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1342
* 20:23 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 20:23 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1342.eqiad.wmnet
* 20:23 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1342.eqiad.wmnet
* 20:22 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1342.eqiad.wmnet
* 20:15 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1027.eqiad.wmnet with OS bookworm
* 19:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:41 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1027.eqiad.wmnet with OS bookworm
* 19:36 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:28 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:26 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:19 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 19:19 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:06 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 19:05 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 19:03 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 18:59 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 18:23 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:06 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1026.eqiad.wmnet with OS bookworm
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 18:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 17:54 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 17:49 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 17:45 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1026.eqiad.wmnet with OS bookworm
* 17:31 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:29 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 17:27 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 17:12 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 17:04 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 16:46 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] (duration: 09m 28s)
* 16:41 urbanecm@deploy1003: migr, urbanecm: Continuing with deployment
* 16:41 urbanecm@deploy1003: migr, urbanecm: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:36 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]]
* 16:28 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:25 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 15:55 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 15:54 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 15:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] (duration: 09m 43s)
* 15:41 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 15:40 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:36 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]]
* 15:17 urbanecm: Delete all running periodic jobs starting with `growthexperiments-refreshlinkrecommendations-*` (to pick up new configuration; [[phab:T392944|T392944]])
* 15:08 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:06 moritzm: installing grub2 bugfix updates from Bookworm point release
* 15:04 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp6008.drmrs.wmnet
* 15:01 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:42 hnowlan: half concurrency for parsoidCachePrewarm RecordLintJob and refreshLinks in jobqueue, eqiad & codfw
* 14:35 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 14:34 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:32 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 14:20 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:07 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2013.codfw.wmnet with OS trixie
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 14:06 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 13:55 btullis@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'.
* 13:53 btullis@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'.
* 13:43 btullis@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 13:42 btullis@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 13:29 moritzm: pruned obsolete Bullseye image dispatch from the docker registry [[phab:T416452|T416452]]
* 13:28 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:26 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b7-eqiad
* 13:25 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 13:24 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:22 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 13:22 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:17 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a4-eqiad
* 13:17 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] (duration: 10m 15s)
* 13:14 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 13:11 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 13:10 sbisson@deploy1003: sbisson: Continuing with deployment
* 13:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:04 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]]
* 13:02 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on ldap-rw[1001,2001].wikimedia.org with reason: work in progress
* 12:49 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 12:48 btullis@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
* 12:46 btullis@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
* 12:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] (duration: 14m 39s)
* 12:41 ladsgroup@deploy1003: tryvix1509, ladsgroup: Continuing with deployment
* 12:35 ladsgroup@deploy1003: tryvix1509, ladsgroup: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:31 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]]
* 12:16 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 12:16 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 11:53 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] (duration: 21m 58s)
* 11:48 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 11:35 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:31 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]]
* 10:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Repooling db2196
* 10:47 moritzm: pruned obsolete Bullseye images nodejs12-slim/nodejs12-devel/nodejs14-slim/nodejs16-slim from the docker registry [[phab:T416452|T416452]]
* 10:43 moritzm: installing Bird security updates
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1260: Repooling after cloning
* 10:09 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Repooling db2196
* 10:07 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 10:06 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 09:55 moritzm: pruned obsolete Bullseye images openjdk-8-jdk/openjdk-8-jre/openjdk-11-jre/openjdk-11-jdk from the docker registry [[phab:T416452|T416452]]
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1260: Repooling after cloning
* 09:52 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 09:52 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 09:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:28 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:27 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:03 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 09:02 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 09:01 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 16 hosts with reason: upgrade ssw1-a1-eqiad
* 08:58 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 22 hosts with reason: upgrade ssw1-a1-eqiad
* 08:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 08:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 08:49 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 08:49 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 08:48 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 08:48 cmooney@cumin1004: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 08:36 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning sanitarium
* 08:30 brouberol@dns1004: END - running authdns-update
* 08:28 moritzm: pruned obsolete Bullseye image golang1.15 from the docker registry [[phab:T416452|T416452]]
* 08:28 brouberol@dns1004: START - running authdns-update
* 08:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1260: Needs to clone another host from this one
* 08:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1260: Needs to clone another host from this one
* 08:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1260.eqiad.wmnet with reason: Cloning sanitarium
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:00 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 07:40 chlod: UTC morning backport window done
* 07:37 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] (duration: 21m 36s)
* 07:32 chlod@deploy1003: chlod, hamishz: Continuing with deployment
* 07:20 chlod@deploy1003: chlod, hamishz: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 45s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1025.eqiad.wmnet with OS bookworm
== 2026-09-08 ==
* 23:51 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:48 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:39 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1313.eqiad.wmnet
* 23:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:25 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 23:19 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:19 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:15 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:05 Amir1: dropped 57 tables on db1260 ([[phab:T437278|T437278]])
* 23:03 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 23:03 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:02 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 22:57 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1313
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1313
* 22:37 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1313
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:37 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:33 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 22:33 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1313
* 22:32 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 22:32 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1313.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1306.eqiad.wmnet
* 22:27 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6008.drmrs.wmnet with OS trixie
* 22:18 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 22:03 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 22:01 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] (duration: 09m 53s)
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 21:58 Amir1: drop links tables from db2210 ([[phab:T437278|T437278]])
* 21:57 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 21:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:56 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:52 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 21:51 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]]
* 21:51 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:48 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:45 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] (duration: 05m 27s)
* 21:43 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Continuing with deployment
* 21:40 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:39 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]]
* 21:35 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1024.eqiad.wmnet with OS bookworm
* 21:33 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp6008.drmrs.wmnet with OS trixie
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1306
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1306
* 21:30 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1306
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:24 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] (duration: 09m 12s)
* 21:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 21:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1306
* 21:19 reedy@deploy1003: reedy: Continuing with deployment
* 21:19 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:19 reedy@deploy1003: reedy: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1306.eqiad.wmnet
* 21:15 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:15 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]]
* {{safesubst:SAL entry|1=21:10 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out}}
* 21:05 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1024.eqiad.wmnet with OS bookworm
* 21:05 sbassett@deploy1003: sbassett: Continuing with deployment
* 21:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 21:04 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=21:03 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out non-http(s) lice}}
* 20:59 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=20:58 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out n}}
* 20:53 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 20:50 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1305.eqiad.wmnet
* 20:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 20:34 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 20:28 sbassett@deploy1003: sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:27 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-http(s) license}}
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1023.eqiad.wmnet with OS bookworm
* {{safesubst:SAL entry|1=20:23 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-}}
* 20:15 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] (duration: 10m 16s)
* 20:13 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:10 aaron@deploy1003: aaron: Continuing with deployment
* 20:09 aaron@deploy1003: aaron: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 20:05 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:05 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]]
* 20:01 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1023.eqiad.wmnet with OS bookworm
* 19:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1305
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1305
* 19:43 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1305
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:43 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:39 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 19:39 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1305
* 19:38 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 19:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1305.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1275.eqiad.wmnet
* 19:23 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 19:17 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 18:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:53 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1275
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1275
* 18:33 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1275
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1275
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 18:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1275.eqiad.wmnet
* 18:18 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 17:43 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy1003.eqiad.wmnet
* 17:36 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet
* 17:36 cdobbins@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-ntp (exit_code=0) rolling restart_daemons on A:dnsbox
* 17:30 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy1003.eqiad.wmnet
* 17:25 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet
* 17:15 swfrench@deploy1003: Finished scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup (duration: 04m 18s)
* 17:11 Amir1: dropping links tables from db1247 (s4 replica) - ([[phab:T437278|T437278]])
* 17:10 swfrench@deploy1003: Started scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup
* 16:51 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] (duration: 10m 19s)
* 16:46 zabe@deploy1003: zabe: Continuing with deployment
* 16:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]]
* 16:29 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 jhancock@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:12 jhancock@cumin2003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 16:08 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2003.codfw.wmnet
* 15:58 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:55 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2003.codfw.wmnet
* 15:54 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:44 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2002.codfw.wmnet
* 15:19 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2002.codfw.wmnet
* 14:55 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:46 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2001.codfw.wmnet
* 14:44 topranks: shutdown et-1/1/5 on cr1-codfw to shift traffic off ssw1-a1-codfw
* 14:43 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: upgrade ssw1-a1-eqiad
* 14:34 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2001.codfw.wmnet
* 14:33 btullis@cumin1003: END (ERROR) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=97) rolling reboot on A:cephosd-codfw
* 14:30 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet
* 14:28 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet
* 14:22 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet
* 14:13 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --olderThan {{Gerrit|1788220800}} --verbose # [[phab:T437158|T437158]]
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet
* 14:12 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet
* 14:03 topranks: drain traffic from ssw1-a1-codfw before JunOS upgrade [[phab:T426197|T426197]]
* 14:02 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet
* 13:58 cgoubert@deploy1003: helmfile [staging-codfw] DONE helmfile.d/services/mw-debug: apply
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet
* 13:57 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet
* 13:56 cgoubert@deploy1003: helmfile [staging-codfw] START helmfile.d/services/mw-debug: apply
* 13:55 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:54 stran@deploy1003: mwscript-k8s job started: foreachwikiindblist checkuser-suggested-investigations extensions/CheckUser/maintenance/populateSiCaseProperties.php # [[phab:T435066|T435066]]
* 13:52 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:51 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet
* 13:50 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:46 cdobbins@cumin1003: START - Cookbook sre.dns.roll-restart-ntp rolling restart_daemons on A:dnsbox
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet
* 13:46 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet
* 13:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2012.codfw.wmnet with OS bookworm
* 13:44 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] (duration: 34m 00s)
* 13:40 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet
* 13:37 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet
* 13:35 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet
* 13:32 stran@deploy1003: stran: Continuing with deployment
* 13:29 stran@deploy1003: stran: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet
* 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:21 moritzm: installing qemu security updates
* 13:18 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet
* 13:16 ayounsi@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'email' for AS: 139628
* 13:15 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 139628
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet
* 13:13 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet
* 13:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 2519
* 13:11 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 2519
* 13:10 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 14593
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]]
* 13:09 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:08 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 14593
* 13:06 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet
* 13:06 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2012.codfw.wmnet with OS bookworm
* 13:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2012.codfw.wmnet
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2012.codfw.wmnet
* 13:05 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:04 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet
* 13:01 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet
* 12:58 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 34655
* 12:58 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 34655
* 12:56 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2012.codfw.wmnet
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'clear' for AS: 35320
* 12:55 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'clear' for AS: 35320
* 12:55 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-codfw
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-codfw
* 12:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2011.codfw.wmnet with OS bookworm
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-by27-esams
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-by27-esams
* 12:54 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-esams
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-esams
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-esams
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-esams
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-eqsin
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-eqsin
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-ulsfo
* 12:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f3-eqiad
* 12:50 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f3-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-c8-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-c8-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:44 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:43 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet
* 12:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-eqiad
* 12:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f1-eqiad
* 12:41 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f1-eqiad
* 12:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-d5-eqiad
* 12:40 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-d5-eqiad
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet
* 12:38 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet
* 12:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:32 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet
* 12:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad
* 12:11 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2012.codfw.wmnet
* 12:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2011.codfw.wmnet with OS bookworm
* 12:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2011.codfw.wmnet
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2011.codfw.wmnet
* 12:09 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:06 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] (duration: 09m 54s)
* 12:01 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2011.codfw.wmnet
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Continuing with deployment
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2010.codfw.wmnet with OS bookworm
* 11:56 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]]
* 11:46 marostegui@dns1004: END - running authdns-update
* 11:44 marostegui@dns1004: START - running authdns-update
* 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:40 Amir1: dropping unneeded tables from x4 - db1260 ([[phab:T437278|T437278]])
* 11:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:23 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2011.codfw.wmnet
* 11:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2010.codfw.wmnet with OS bookworm
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2010.codfw.wmnet
* 11:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2010.codfw.wmnet
* 11:20 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2010.codfw.wmnet
* 11:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2009.codfw.wmnet with OS bookworm
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:43 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] (duration: 10m 57s)
* 10:39 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:38 samtar@deploy1003: samtar: Continuing with deployment
* 10:37 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:37 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:36 samtar@deploy1003: samtar: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:34 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:32 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]]
* 10:30 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2009.codfw.wmnet with OS bookworm
* 10:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2009.codfw.wmnet
* 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:18 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2009.codfw.wmnet
* 10:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2008.codfw.wmnet with OS bookworm
* 10:07 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:05 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:01 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:45 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] (duration: 13m 15s)
* 09:45 ayounsi@dns1004: END - running authdns-update
* 09:44 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:43 ayounsi@dns1004: START - running authdns-update
* 09:39 zabe@deploy1003: zabe: Continuing with deployment
* 09:37 zabe@deploy1003: zabe: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2008.codfw.wmnet with OS bookworm
* 09:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:32 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]]
* 09:32 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2008.codfw.wmnet
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2008.codfw.wmnet
* 09:31 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:29 XioNoX: remove GRE tunnels eqiad-drmrs eqdfw-ulsfo
* 09:23 moritzm: installing rsync security updates
* 09:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2007.codfw.wmnet with OS bookworm
* 09:22 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2008.codfw.wmnet
* 09:11 marostegui@cumin1003: dbctl commit (dc=all): 'Make x4 and s4 RW again [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96393 and previous config saved to /var/cache/conftool/dbconfig/20260908-091121-marostegui.json
* 09:07 marostegui@cumin1003: dbctl commit (dc=all): 'Remove old s4 masters from x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96392 and previous config saved to /var/cache/conftool/dbconfig/20260908-090749-marostegui.json
* 09:05 marostegui@cumin1003: dbctl commit (dc=all): 'Set x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96391 and previous config saved to /var/cache/conftool/dbconfig/20260908-090517-marostegui.json
* 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Set s4 commons to read-only for maintenance [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96389 and previous config saved to /var/cache/conftool/dbconfig/20260908-090228-marostegui.json
* 09:02 marostegui: Starting x4 split from s4, RO time on commons needed [[phab:T404715|T404715]]
* 09:00 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 08:58 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2008.codfw.wmnet
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:43 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 32 hosts with reason: x4 split
* 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2007.codfw.wmnet with OS bookworm
* 08:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:37 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:37 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2007.codfw.wmnet
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2007.codfw.wmnet
* 08:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:34 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:29 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:26 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2007.codfw.wmnet
* 08:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2006.codfw.wmnet with OS bookworm
* 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:50 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2007.codfw.wmnet
* 07:43 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2006.codfw.wmnet with OS bookworm
* 07:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2006.codfw.wmnet
* 07:37 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:30 denisse: Add grafana-plugins 0.15 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 07:29 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:27 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 07:27 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:27 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:22 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 07:22 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 07:18 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:14 jmm@dns1004: END - running authdns-update
* 07:13 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96388 and previous config saved to /var/cache/conftool/dbconfig/20260908-071308-marostegui.json
* 07:12 jmm@dns1004: START - running authdns-update
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96387 and previous config saved to /var/cache/conftool/dbconfig/20260908-071216-marostegui.json
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96386 and previous config saved to /var/cache/conftool/dbconfig/20260908-071159-marostegui.json
* 05:07 denisse: Add grafana-plugins 0.10 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.16 (duration: 02m 27s)
* 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]] (duration: 36m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 41s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-07 ==
* 21:52 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] (duration: 11m 00s)
* 21:47 zabe@deploy1003: zabe: Continuing with deployment
* 21:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]]
* 21:37 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] (duration: 09m 34s)
* 21:33 zabe@deploy1003: zabe: Continuing with deployment
* 21:32 zabe@deploy1003: zabe: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:28 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]]
* 21:03 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] (duration: 10m 27s)
* 20:58 zabe@deploy1003: zabe: Continuing with deployment
* 20:57 zabe@deploy1003: zabe: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:52 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]]
* 20:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set weight of db1261 to zero in s4 ([[phab:T437108|T437108]])', diff saved to https://phabricator.wikimedia.org/P96385 and previous config saved to /var/cache/conftool/dbconfig/20260907-203804-ladsgroup.json
* 20:23 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] (duration: 09m 28s)
* 20:19 zabe@deploy1003: zabe: Continuing with deployment
* 20:18 zabe@deploy1003: zabe: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:14 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]]
* 20:12 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] (duration: 10m 06s)
* 20:07 zabe@deploy1003: zabe, daimona: Continuing with deployment
* 20:06 zabe@deploy1003: zabe, daimona: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:02 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]]
* 19:59 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] (duration: 11m 27s)
* 19:55 zabe@deploy1003: zabe: Continuing with deployment
* 19:52 zabe@deploy1003: zabe: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:48 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]]
* 19:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] (duration: 11m 40s)
* 19:26 zabe@deploy1003: zabe: Continuing with deployment
* 19:23 zabe@deploy1003: zabe: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:19 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]]
* 19:07 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] (duration: 14m 17s)
* 19:00 zabe@deploy1003: zabe: Continuing with deployment
* 18:57 zabe@deploy1003: zabe: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:53 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]]
* 18:33 zabe@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 18:32 zabe@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 18:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] (duration: 09m 12s)
* 18:27 zabe@deploy1003: zabe: Continuing with deployment
* 18:26 zabe@deploy1003: zabe: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:22 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]]
* 16:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2005.codfw.wmnet with OS bookworm
* 16:01 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] (duration: 10m 22s)
* 15:59 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 15:57 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 15:56 zabe@deploy1003: zabe: Continuing with deployment
* 15:55 zabe@deploy1003: zabe: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:51 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]]
* 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:47 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 15:46 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 15:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:44 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 15:44 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 15:27 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2005.codfw.wmnet with OS bookworm
* 15:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2005.codfw.wmnet
* 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:11 moritzm: installing rsync security updates
* 15:04 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1341.eqiad.wmnet
* 15:03 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2004.codfw.wmnet with OS bookworm
* 15:00 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:58 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 14:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:40 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 14:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 14:35 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:32 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:29 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:25 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 14:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1228: Repooling db1228 into s4
* 14:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2004.codfw.wmnet with OS bookworm
* 14:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Repooling after cloning
* 14:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:19 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2004.codfw.wmnet
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2004.codfw.wmnet
* 14:18 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:14 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1074.eqiad.wmnet
* 14:14 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:13 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:13 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* {{safesubst:SAL entry|1=14:11 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mul}}
* 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2004.codfw.wmnet
* 14:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1074.eqiad.wmnet
* 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1073.eqiad.wmnet
* 14:07 krinkle@deploy1003: krinkle: Continuing with deployment
* {{safesubst:SAL entry|1=14:04 krinkle@deploy1003: krinkle: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with multiple properties}}
* 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1073.eqiad.wmnet
* 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1072.eqiad.wmnet
* 14:01 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1341
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1341
* 14:01 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 14:00 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1341
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* 13:59 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* {{safesubst:SAL entry|1=13:59 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mult}}
* 13:59 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2004.codfw.wmnet
* 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2003.codfw.wmnet with OS bookworm
* 13:55 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 13:55 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1072.eqiad.wmnet
* 13:55 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudvirt1067.eqiad.wmnet
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1341
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 13:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1341.eqiad.wmnet
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2008.wikimedia.org
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2008.wikimedia.org with OS trixie
* 13:39 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1067.eqiad.wmnet
* 13:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1066.eqiad.wmnet
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:37 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1228: Repooling db1228 into s4
* 13:36 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Repooling after cloning
* 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1066.eqiad.wmnet
* 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1065.eqiad.wmnet
* 13:28 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:27 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1065.eqiad.wmnet
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:25 moritzm: installing openssh security updates
* 13:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:24 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1340.eqiad.wmnet
* 13:23 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] (duration: 10m 06s)
* 13:17 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2003.codfw.wmnet with OS bookworm
* 13:17 stran@deploy1003: stran: Continuing with deployment
* 13:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:16 stran@deploy1003: stran: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:16 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:15 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 13:12 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]]
* 13:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2003.codfw.wmnet
* 13:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2003.codfw.wmnet
* 13:03 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2008.wikimedia.org with OS trixie
* 13:03 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:02 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:02 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:01 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:00 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2003.codfw.wmnet
* 12:58 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:54 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 12:54 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2008.wikimedia.org
* 12:47 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2003.codfw.wmnet
* 12:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2002.codfw.wmnet with OS bookworm
* 12:29 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 12:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2007.wikimedia.org
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2007.wikimedia.org with OS trixie
* 12:14 elukey: moved most of the Docker Registry's prefixes to a new internal S3 backend. For any docker pull failure that worked in the past, please ping me or drop a note in [[phab:T435499|T435499]] or contact the oncall SREs
* 12:07 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2002.codfw.wmnet with OS bookworm
* 12:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:02 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2002.codfw.wmnet
* 11:54 jmm@dns1004: END - running authdns-update
* 11:52 jmm@dns1004: START - running authdns-update
* 11:46 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2002.codfw.wmnet
* 11:46 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2007.wikimedia.org with OS trixie
* 11:46 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:46 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:41 moritzm: installing bash updates from bookworm point release
* 11:39 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:39 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2007.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts ldap-replica1006.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:35 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:32 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2001.codfw.wmnet with OS bookworm
* 11:28 jmm@cumin1004: START - Cookbook sre.hosts.decommission for hosts ldap-replica1006.wikimedia.org
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 11:18 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] (duration: 14m 08s)
* 11:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 11:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet
* 11:11 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:11 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:10 zabe@deploy1003: zabe: Continuing with deployment
* 11:10 zabe@deploy1003: zabe: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 11:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet
* 11:07 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet
* 11:07 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet
* 11:06 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:06 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:04 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]]
* 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 10:58 btullis@deploy1003: Finished scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]] (duration: 35m 20s)
* 10:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:53 jmm@dns1004: END - running authdns-update
* 10:51 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:50 jmm@dns1004: START - running authdns-update
* 10:47 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=frwiki # [[phab:T436659|T436659]]
* 10:40 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=hrwiki # [[phab:T436659|T436659]]
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:39 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 10:37 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 10:33 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 10:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2001.codfw.wmnet with OS bookworm
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1340
* 10:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1340.eqiad.wmnet
* 10:26 btullis@deploy1003: Started scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]]
* 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2001.codfw.wmnet
* 10:18 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2001.codfw.wmnet
* 10:13 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw
* 10:12 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:10 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:09 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:08 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:07 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:07 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:02 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:00 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:59 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:57 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/services/miscweb: apply
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/services/miscweb: apply
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:52 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:52 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] (duration: 10m 11s)
* 09:51 aokoth@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 09:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 09:49 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 09:49 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 09:48 aokoth@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 09:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:45 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:44 aokoth@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 09:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]]
* 09:41 aokoth@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 09:38 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Repooling after switchover
* 09:23 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw
* 09:15 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] (duration: 20m 12s)
* 09:12 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 139009
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1005.wikimedia.org
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1005.wikimedia.org with OS trixie
* 09:10 moritzm: rebuild software RAID following disk replacement [[phab:T437036|T437036]]
* 09:10 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 139009
* 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1022.eqiad.wmnet with OS bookworm
* 09:08 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 09:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:03 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet
* 09:02 moritzm: installing giflib security updates
* 09:01 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:59 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 08:57 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 08:56 jmm@cumin1004: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet
* 08:56 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:54 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]]
* 08:52 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:52 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:49 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 08:48 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:42 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:41 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:40 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:40 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96355 and previous config saved to /var/cache/conftool/dbconfig/20260907-083904-marostegui.json
* 08:38 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1005.wikimedia.org with OS trixie
* 08:38 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2157 to s5 primary [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96354 and previous config saved to /var/cache/conftool/dbconfig/20260907-083825-marostegui.json
* 08:38 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 marostegui: Starting s5 codfw failover from db2213 to db2157 - [[phab:T437188|T437188]]
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2157 with weight 0 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96353 and previous config saved to /var/cache/conftool/dbconfig/20260907-083448-marostegui.json
* 08:34 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s5 [[phab:T437188|T437188]]
* 08:28 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:28 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1005.wikimedia.org
* 08:22 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:20 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1022.eqiad.wmnet with OS bookworm
* 08:03 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:02 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 08:02 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:00 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 07:57 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 07:56 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1241.eqiad.wmnet with reason: Cloning
* 07:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1241: Cloning
* 07:55 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1241: Cloning
* 07:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Cloning
* 07:54 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Cloning
* 07:51 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 07:50 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 07:47 kartik@deploy1003: Finished scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] (duration: 41m 51s)
* 07:46 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 07:45 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 07:34 kartik@deploy1003: abi, kartik: Continuing with deployment
* 07:23 kartik@deploy1003: abi, kartik: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:05 kartik@deploy1003: Started scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]]
* 06:14 moritzm: installing Chromium security updates
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 08m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-06 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 25s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-05 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 26s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-04 ==
* 22:07 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 21:48 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:42 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:30 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:42 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:40 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 19:27 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:19 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:13 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:12 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host sretest2013
* 19:11 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host sretest2013
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:11 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:07 jhancock@cumin1003: START - Cookbook sre.dns.netbox
* 18:18 inflatador: bking@clouddumps100[12] `systemctl reset-failed` to quash alerts until https://w.wiki/UBje . The systemd timer should try again tomorrow
* 17:27 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b8-eqiad
* 17:27 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b8-eqiad
* 16:37 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:33 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-b7-eqiad
* 16:33 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 16:05 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b6-eqiad
* 16:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b6-eqiad
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1339.eqiad.wmnet
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 15:50 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:49 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:47 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:43 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 15:10 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b5-eqiad
* 15:09 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b5-eqiad
* 14:46 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1045.eqiad.wmnet
* 14:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1003.eqiad.wmnet with OS trixie
* 14:38 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 14:37 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b4-eqiad
* 14:37 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b4-eqiad
* 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339
* 14:32 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339
* 14:31 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 14:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet
* 14:26 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:24 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:21 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b3-eqiad
* 14:21 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b3-eqiad
* 14:17 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:17 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:04 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1003.eqiad.wmnet with OS trixie
* 13:54 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b2-eqiad
* 13:53 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b2-eqiad
* 13:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1002.eqiad.wmnet with OS trixie
* 13:18 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a4-eqiad
* 13:12 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:12 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b1-eqiad
* 13:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b1-eqiad
* 12:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1002.eqiad.wmnet with OS trixie
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow3004.esams.wmnet with OS trixie
* 12:33 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:28 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 12:15 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 12:11 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a7-eqiad
* 12:11 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a7-eqiad
* 12:01 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow3004.esams.wmnet with OS trixie
* 11:47 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a6-eqiad
* 11:47 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a6-eqiad
* 11:36 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki2003.codfw.wmnet
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki2003.codfw.wmnet with OS trixie
* 11:19 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:13 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a5-eqiad
* 11:06 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a5-eqiad
* 10:52 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki2003.codfw.wmnet with OS trixie
* 10:50 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:50 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:44 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 10:44 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki2003.codfw.wmnet
* 10:29 btullis@deploy1003: Finished scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli (duration: 41m 14s)
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 10:00 marostegui@cumin1003: Removing db1182 from zarcillo [[phab:T434869|T434869]]
* 10:00 marostegui@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db1182.eqiad.wmnet
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:57 btullis@deploy1003: Started scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli
* 09:53 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:53 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.decommission (exit_code=1)
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1182.eqiad.wmnet
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:50 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:46 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:45 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1182 from dbctl [[phab:T434869|T434869]]', diff saved to https://phabricator.wikimedia.org/P96346 and previous config saved to /var/cache/conftool/dbconfig/20260904-094527-marostegui.json
* 09:41 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki1003.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki1003.eqiad.wmnet with OS trixie
* 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1182: Decommissioning
* 09:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1182: Decommissioning
* 09:23 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 09:17 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow5003.eqsin.wmnet with OS trixie
* 09:02 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki1003.eqiad.wmnet with OS trixie
* 09:00 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 09:00 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:55 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:55 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki1003.eqiad.wmnet
* 08:54 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:48 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:45 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:40 btullis@deploy1003: Finished scap sync-world: Trying again for [[phab:T436913|T436913]] (duration: 34m 26s)
* 08:35 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2003.codfw.wmnet
* 08:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw2001.wikimedia.org with OS trixie
* 08:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4 days, 0:00:00 on db2196.codfw.wmnet with reason: Host crashed
* 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2003.codfw.wmnet
* 08:24 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:20 elukey@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:12 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:08 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:07 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 08:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:04 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 07:57 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow5003.eqsin.wmnet with OS trixie
* 07:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:55 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw2001.wikimedia.org with OS trixie
* 07:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:51 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:50 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie
* 07:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2196: down
* 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2196: down
* 06:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:52 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:40 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:28 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:21 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 06:20 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 39s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 01:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:02 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:59 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:46 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 00:44 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:39 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:23 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
== 2026-09-03 ==
* 21:46 tsev@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T435363|T435363]]
* 21:03 eevans@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:55 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] (duration: 12m 24s)
* 20:52 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:52 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1021.eqiad.wmnet with OS bookworm
* 20:50 arlolra@deploy1003: arlolra, tsev: Continuing with deployment
* 20:46 arlolra@deploy1003: arlolra, tsev: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:44 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1047.eqiad.wmnet
* 20:42 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]]
* 20:41 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a3-eqiad
* 20:40 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a3-eqiad
* 20:39 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] (duration: 10m 23s)
* 20:34 arlolra@deploy1003: arlolra, jgiannelos: Continuing with deployment
* 20:33 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1047.eqiad.wmnet
* 20:33 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:32 arlolra@deploy1003: arlolra, jgiannelos: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:31 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1046.eqiad.wmnet
* 20:28 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]]
* 20:23 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] (duration: 13m 41s)
* 20:20 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1046.eqiad.wmnet
* 20:16 catrope@deploy1003: catrope: Continuing with deployment
* 20:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1021.eqiad.wmnet with OS bookworm
* 20:13 catrope@deploy1003: catrope: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1021.eqiad.wmnet
* 20:09 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]]
* 20:00 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1021.eqiad.wmnet
* 19:49 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 19:19 swfrench@deploy1003: Finished scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]] (duration: 02m 59s)
* 19:16 swfrench@deploy1003: Started scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]]
* 19:01 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:01 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 18:57 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:57 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:34 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:20 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 17:55 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:55 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:53 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:52 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:46 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 17:46 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 17:45 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 17:44 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 17:44 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:40 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:39 ryankemper: [WDQS] Service looks healthy again, CPU load and thread count have dropped considerably over the last hour
* 17:39 ryankemper: [[phab:T421642|T421642]] [WDQS] requestctl changes: `2026-09-03 16:23-17:33` UTC: added hard-deny pair `cache-text/wdqs_futile_sparql_sep_2026_deny(+_bots)`; extended pattern `ua/wdqs_heavy_sparql_bots_2026` and added default-scope twin `wdqs_heavy_sparql_bots_jul_2026_ratelimit_default`; added ipblock `abuse/wdqs_sparql_scanners_sep_2026` + throttle `wdqs_sparql_scanners_sep_2026_ratelimit` (needed manual `requestctl update-provenance-map`)
* 17:37 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:37 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:31 andrew@cumin2003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudcephosd1045.eqiad.wmnet
* 17:30 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:30 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:28 dancy@deploy1003: Installation of scap version "4.289.0" completed for 3 hosts
* 17:26 dancy@deploy1003: Installing scap version "4.289.0" for 3 host(s)
* 17:24 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a2-eqiad
* 17:24 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:22 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:22 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a2-eqiad
* 17:14 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 17:10 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:10 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:09 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 17:01 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:59 andrew@cumin2003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 16:58 dancy: Running scap clean-images on deploy1003
* 16:52 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:50 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:39 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:14 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs2021.codfw.wmnet,service=wdqs-main
* 16:14 ryankemper: [[phab:T430880|T430880]] Stumbled across `wdqs2021` listed as inactive, looks like it was never fully re-pooled after a data xfer. Pooled.
* 16:12 btullis@deploy1003: Finished deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2] (duration: 00m 38s)
* 16:12 btullis@deploy1003: Started deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2]
* 16:12 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:07 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs101[1-4].eqiad.wmnet
* 16:03 btullis@deploy1003: Started scap sync-world: Rebuilding to pick up new version of dump scripts in mediawiki-cli for [[phab:T436913|T436913]]
* 16:01 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]] (duration: 09m 29s)
* 15:51 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=urldownloader[12]00[56].wikimedia.org [reason: depooling urldownloader trixie nodes]
* 15:51 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]]
* 15:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:29 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:24 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:24 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:18 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:15 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:09 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:05 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 15:04 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:02 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 15:00 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader
* 14:58 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw
* 14:53 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-ulsfo
* 14:49 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 14:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2002.codfw.wmnet
* 14:45 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:39 sukhe: sudo cumin "A:cp-text" "run-puppet-agent --enable 'merging CR 1334855'": [[phab:T425441|T425441]]
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2002.codfw.wmnet
* 14:32 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:30 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:29 sukhe: sudo cumin "A:cp-text" "disable-puppet 'merging CR 1334855'" [[phab:T425441|T425441]]
* 14:27 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-ulsfo
* 14:22 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:21 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:19 arnaudb@dns1006: END - running authdns-update
* 14:18 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1074
* 14:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1074
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-codfw
* 14:17 arnaudb@dns1006: START - running authdns-update
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:16 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:16 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1073
* 14:15 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1073
* 14:12 ayounsi@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host netflow2004.codfw.wmnet with OS trixie
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:10 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:08 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:08 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:07 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] (duration: 09m 36s)
* 14:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:05 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:03 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 14:03 samtar@deploy1003: btullis, samtar: Continuing with deployment
* 14:02 samtar@deploy1003: btullis, samtar: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:01 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:58 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]]
* 13:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:56 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:55 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:52 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqiad and A:durum
* 13:52 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-codfw
* 13:51 moritzm: installing sqlite3 security updates
* 13:51 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:51 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqiad and A:durum
* 13:49 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-codfw and A:durum
* 13:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:47 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-codfw and A:durum
* 13:47 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-esams
* 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:44 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:43 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:43 ayounsi@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:42 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] (duration: 13m 50s)
* 13:41 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:40 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1072
* 13:40 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:39 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-esams and A:durum
* 13:38 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1072
* 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:38 samtar@deploy1003: samtar, thiemowmde: Continuing with deployment
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-esams and A:durum
* 13:37 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:33 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:33 samtar@deploy1003: samtar, thiemowmde: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:32 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqsin and A:durum
* 13:31 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqsin and A:durum
* 13:28 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]]
* 13:28 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 13:27 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:24 ayounsi@cumin1004: START - Cookbook sre.hosts.reimage for host netflow2004.codfw.wmnet with OS trixie
* 13:24 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:22 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 13:22 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-esams
* 13:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:15 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:15 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 moritzm: installing bash updates from trixie point release
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 13:14 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1067
* 13:13 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1067
* 13:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2901: Test
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:09 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:09 moritzm: installing libxslt bugfix updates from Trixie point release
* 13:08 jelto@dns1004: END - running authdns-update
* 13:06 jelto@dns1004: START - running authdns-update
* 13:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 13:04 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:04 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 13:00 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:59 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1066
* 12:59 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1066
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:58 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:54 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 12:53 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2901: Test
* 12:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2901: Test
* 12:52 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'sync'.
* 12:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 12:50 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2901: Test
* 12:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:50 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 12:42 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:41 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-magru and A:durum
* 12:38 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-magru and A:durum
* 12:34 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 12:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1282: Pooling db1282 into s6
* 12:31 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 12:25 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:21 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:19 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:15 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_ulsfo
* 12:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1020.eqiad.wmnet with OS bookworm
* 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: db2207 repool
* 12:07 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_ulsfo
* 12:04 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 11:58 kart_: cxserver: Use urldownloader LVS endpoint ([[phab:T429175|T429175]])
* 11:57 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply
* 11:56 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply
* 11:56 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_magru
* 11:55 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply
* 11:55 moritzm: installing rsync security updates
* 11:55 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply
* 11:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:52 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply
* 11:51 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply
* 11:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1282: Pooling db1282 into s6
* 11:45 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1282 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96328 and previous config saved to /var/cache/conftool/dbconfig/20260903-114526-marostegui.json
* 11:43 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqiad
* 11:35 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqiad
* 11:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1020.eqiad.wmnet with OS bookworm
* 11:26 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 11:24 cgoubert@deploy1003: Finished scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter (duration: 12m 01s)
* 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: db2207 repool
* 11:22 cgoubert@deploy1003: cgoubert: Continuing with deployment
* 11:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqsin
* 11:17 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_esams
* 11:15 cgoubert@deploy1003: cgoubert: mediawiki: enable forward of fatal metrics to statsd exporter synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:14 cgoubert@deploy1003: Started scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter
* 11:10 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_esams
* 11:09 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 11:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1019.eqiad.wmnet with OS bookworm
* 11:01 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_drmrs
* 10:59 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_codfw
* 10:52 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_codfw
* 10:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:41 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 10:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 10:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow2005.codfw.wmnet
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2005.codfw.wmnet with OS trixie
* 10:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:17 btullis@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync
* 10:17 btullis@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync
* 10:16 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1019.eqiad.wmnet with OS bookworm
* 10:15 btullis@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: sync
* 10:15 btullis@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: sync
* 10:12 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s8
* 10:11 marostegui: Move s8 sanitarium from db1167 to db1281 [[phab:T434778|T434778]]
* 10:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 10:03 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 09:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:55 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2003.codfw.wmnet with OS trixie
* 09:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:43 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow2005.codfw.wmnet with OS trixie
* 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:42 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1018.eqiad.wmnet with OS bookworm
* 09:41 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_ulsfo
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:39 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2003.codfw.wmnet with reason: host reimage
* 09:39 hnowlan: fixed currently oncall pane in klaxon
* 09:38 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:38 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434751|T434751]]
* 09:37 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:37 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow2005.codfw.wmnet
* 09:35 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434775|T434775]]
* 09:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s7
* 09:34 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:33 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_ulsfo
* 09:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqiad
* 09:30 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] (duration: 09m 30s)
* 09:27 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:25 zabe@deploy1003: zabe: Continuing with deployment
* 09:25 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqiad
* 09:25 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 09:25 zabe@deploy1003: zabe: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1174 from dbctl [[phab:T436904|T436904]]', diff saved to https://phabricator.wikimedia.org/P96323 and previous config saved to /var/cache/conftool/dbconfig/20260903-092448-marostegui.json
* 09:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:21 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]]
* 09:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqsin
* 09:17 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:15 topranks: put traffic on Lumen codfw<->eqiad link as it is stable [[phab:T435810|T435810]]
* 09:14 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_esams
* 09:09 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:06 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_esams
* 09:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 09:03 marostegui: Move s6 sanitarium from db1165 to db1279 [[phab:T434775|T434775]]
* 09:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1018.eqiad.wmnet with OS bookworm
* 08:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s6
* 08:57 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_drmrs
* 08:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_codfw
* 08:55 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw
* 08:49 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_codfw
* 08:49 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:45 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 08:45 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:42 marostegui: Move s5 sanitarium from db1161 to db1275 [[phab:T434776|T434776]]
* 08:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 24 hosts with reason: Changing sanitarium master in s5
* 08:38 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_magru
* 08:37 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:37 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server (duration: 01m 21s)
* 08:37 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:36 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server
* 08:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:33 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server (duration: 01m 28s)
* 08:32 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server
* 08:27 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 08:11 moritzm: uploaded wmf-laptop 1.0.7 to apt.wikimedia.org
* 08:03 marostegui: Move s2 sanitarium from db1156 to db1271 [[phab:T434287|T434287]]
* 07:59 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2001.codfw.wmnet
* 07:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 23 hosts with reason: Changing sanitarium master in s2
* 07:49 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2001.codfw.wmnet
* 07:39 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:29 chlod: UTC morning backport window done
* 07:27 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] (duration: 11m 54s)
* 07:22 chlod@deploy1003: chlod, tryvix1509: Continuing with deployment
* 07:22 XioNoX: push pfw policies - [[phab:T436729|T436729]]
* 07:20 chlod@deploy1003: chlod, tryvix1509: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]]
* 07:15 marostegui: Power off db1228 for maintenance
* 07:13 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1228.eqiad.wmnet with reason: Onsite maintenance
* 07:01 arnaudb@dns1006: END - running authdns-update
* 06:58 arnaudb@dns1006: START - running authdns-update
* 06:54 jmm@cumin2003: END (PASS) - Cookbook sre.wdqs.restart-nginx-envoy (exit_code=0) rolling restart_daemons on A:wcqs-public
* 06:52 jmm@cumin2003: START - Cookbook sre.wdqs.restart-nginx-envoy rolling restart_daemons on A:wcqs-public
* 06:46 moritzm: installing libxml2 security updates
* 06:27 hashar: Upgrading CI Jenkins on contint1003 # [[phab:T436812|T436812]]
* 06:11 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet
* 06:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts2002.codfw.wmnet
* 06:00 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet
* 05:56 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts2002.codfw.wmnet
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 48s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:16 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1017.eqiad.wmnet with OS bookworm
== 2026-09-02 ==
* 23:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] (duration: 10m 21s)
* 23:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:50 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 23:49 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]]
* 23:38 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:38 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host aqs1017.eqiad.wmnet with OS bookworm
* 22:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] (duration: 11m 03s)
* 22:32 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 22:30 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:26 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]]
* 22:05 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] (duration: 14m 14s)
* 21:59 krinkle@deploy1003: krinkle: Continuing with deployment
* 21:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:50 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]]
* 21:44 inflatador: bking@apt1002 sudo -E private_reprepro --ignore=wrongdistribution -C matomo_plugins include bookworm-wikimedia-private matomo-plugin-customreports_5.5.0-1_amd64.changes [[phab:T431608|T431608]]
* 21:40 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] (duration: 09m 48s)
* 21:35 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:34 jforrester@deploy1003: jforrester: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:34 inflatador: bking@apt1002 sudo -E reprepro -C main include bookworm-wikimedia matomo-plugin-marketingcampaignsreporting_5.2.2-3_amd64.changes [[phab:T431608|T431608]]
* 21:30 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]]
* 21:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 21:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:03 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 20:50 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 20:27 dancy@deploy1003: Finished scap sync-world: testing (duration: 09m 21s)
* 20:18 dancy@deploy1003: Started scap sync-world: testing
* 20:18 dancy@deploy1003: Installation of scap version "4.288.0" completed for 3 hosts
* 20:16 dancy@deploy1003: Installing scap version "4.288.0" for 3 host(s)
* 19:57 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] (duration: 64m 27s)
* 19:55 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 18:57 jforrester@deploy1003: jforrester: Continuing with deployment
* 18:57 jforrester@deploy1003: jforrester: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:55 swfrench-wmf: deleted pods coredns-85b4f68d95-pk5sn coredns-85b4f68d95-22ddb coredns-85b4f68d95-49k5p in eqiad due to intermittent upstream resolution health check failures correlated with high DNS resolution latency
* 18:53 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]]
* 18:31 sukhe@dns1004: END - running authdns-update
* 18:28 sukhe@dns1004: START - running authdns-update
* 18:26 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:26 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:17 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:16 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 18:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:14 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:14 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqiad
* 18:14 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:49 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqiad
* 17:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:45 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqsin
* 17:38 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:38 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:20 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqsin
* 17:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 17:00 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 16:59 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:48 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 16:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:41 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 16:41 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 16:41 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:38 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:34 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 16:29 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_drmrs
* 16:24 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 16:14 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_drmrs
* 16:11 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-unlock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 16:10 root@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]] (duration: 48m 09s)
* 16:10 root@deploy1003: Forcefully removing global lock: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 16:10 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-unlock-scap for datacenter switchover from codfw to eqiad
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:59 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:58 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters for datacenter switchover from codfw to eqiad
* 15:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 15:58 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-restore-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-restore-ttl for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-start-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 root@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-start-maintenance for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-jobrunner: sync
* 15:56 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-jobrunner: sync
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.07-set-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period ends at: 2026-09-02 15:56:13.434320
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.07-set-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.03-set-db-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.03-set-db-readonly for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.02-set-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period starts at: 2026-09-02 15:53:47.690918
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.02-set-readonly for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.01-stop-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.01-stop-maintenance for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-reduce-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:47 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-reduce-ttl for datacenter switchover from codfw to eqiad
* 15:46 slyngshede@cumin1003: END (ERROR) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=97) for datacenter switchover from codfw to eqiad
* 15:45 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 15:44 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 15:42 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service
* 15:39 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:38 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:31 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:28 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:27 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 15:27 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from codfw to eqiad
* 15:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-lock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:22 root@deploy1003: Locking from deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-lock-scap for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks for datacenter switchover from codfw to eqiad
* 15:17 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service
* 15:15 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_magru
* 15:11 moritzm: import jenkins 2.568.3 to thirdparty/jenkins for trixie-wikimedia
* 14:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 14:44 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_magru
* 14:32 moritzm: installing pdns-recursor security updates
* 14:27 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:27 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:20 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:15 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 14:12 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 14:09 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:09 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:09 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:08 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:06 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:44 moritzm: bounce tcpircbot-logmsgbot/tcpircbot-logmsgbot_cloud on alert1002 to allow cumin1004 [[phab:T427897|T427897]]
* 13:36 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] (duration: 09m 52s)
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Continuing with deployment
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:26 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]]
* 13:25 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 13:24 moritzm: installing wireshark security updates
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 13:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:04 moritzm: import librsvg 2.60.0+dfsg-1+wmf13u1 to component/thumbor for trixie-wikimedia [[phab:T436505|T436505]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 12:44 atsuko@dns1004: END - running authdns-update
* 12:41 atsuko@dns1004: START - running authdns-update
* 12:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:35 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] (duration: 12m 50s)
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:24 dreamyjazz@deploy1003: dreamyjazz, btullis: Continuing with deployment
* 12:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:22 dreamyjazz@deploy1003: dreamyjazz, btullis: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:20 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:17 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]]
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 11:31 marostegui@cumin1003: Removing db1172 from zarcillo [[phab:T436763|T436763]]
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1172.eqiad.wmnet
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:30 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:26 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:26 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:25 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 11:25 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:24 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:20 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1172.eqiad.wmnet
* 11:20 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 11:12 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:11 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:10 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:09 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:05 slyngshede@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.*
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 11:05 slyngshede@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 11:03 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply
* 11:03 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
* 10:52 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:49 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:48 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Pool back db1242
* 10:45 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow4003.ulsfo.wmnet with OS trixie
* 10:31 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1172 from dbctl [[phab:T436763|T436763]]', diff saved to https://phabricator.wikimedia.org/P96318 and previous config saved to /var/cache/conftool/dbconfig/20260902-103152-marostegui.json
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 blake@deploy1003: Finished scap sync-world: non-build deployment for [[phab:T417800|T417800]] (duration: 05m 37s)
* 10:06 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:05 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:04 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:03 blake@deploy1003: Started scap sync-world: non-build deployment for [[phab:T417800|T417800]]
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 10:00 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:00 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:59 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:58 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:57 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:56 jmm@dns1004: END - running authdns-update
* 09:56 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:55 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:54 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:53 jmm@dns1004: START - running authdns-update
* 09:51 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:47 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1228 to dbctl [[phab:T435892|T435892]]', diff saved to https://phabricator.wikimedia.org/P96313 and previous config saved to /var/cache/conftool/dbconfig/20260902-094713-marostegui.json
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:42 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow4003.ulsfo.wmnet with OS trixie
* 09:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow7002.magru.wmnet with OS trixie
* 09:31 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:30 moritzm: installing openjdk-21 security updates
* 09:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:22 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:17 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:17 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:16 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:10 moritzm: installing openjdk-8 security updates
* 09:09 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:08 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:46 tappof: bump space for prometheus k8s-dse in eqiad
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:39 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow7002.magru.wmnet with OS trixie
* 08:36 moritzm: installing libgraphite2 security updates
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:23 Msz2001: UTC morning backport window done
* 08:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] (duration: 14m 36s)
* 08:19 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:19 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet
* 08:14 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:14 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet
* 08:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet
* 08:08 fabfur@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cp5022.eqsin.wmnet with reason: investigating
* 08:07 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]]
* 08:07 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp5022.*
* 08:07 fabfur: depooling and silencing cp5022 ([[phab:T414411|T414411]])
* 08:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet
* 08:03 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:03 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* {{safesubst:SAL entry|1=08:03 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)}}
* 07:49 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 07:49 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]] synced to the
* 07:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1004.eqiad.wmnet
* 07:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1004.eqiad.wmnet
* 07:30 jmm@dns1004: END - running authdns-update
* {{safesubst:SAL entry|1=07:27 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]}}
* 07:27 jmm@dns1004: START - running authdns-update
* 07:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] (duration: 16m 04s)
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1242: Cloning db1228
* 07:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1242: Cloning db1228
* 07:18 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1228,1242].eqiad.wmnet with reason: db1242 needs to clone db1228
* 07:17 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Continuing with deployment
* 07:12 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1228.eqiad.wmnet with OS trixie
* 07:10 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be veri
* 07:06 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]]
* 06:43 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: build: Updating npm dependencies (duration: 00m 13s)
* 06:43 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: build: Updating npm dependencies
* 06:39 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:32 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:23 slyngshede@dns1004: END - running authdns-update
* 06:21 marostegui: Drop cu* tables from s3 bswiktionary [[phab:T435965|T435965]]
* 06:20 slyngshede@dns1004: START - running authdns-update
* 06:18 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1228.eqiad.wmnet with OS trixie
* 06:13 XioNoX: re-enable magru cr1/asw1-b3 link - [[phab:T436675|T436675]]
* 05:06 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] (duration: 04m 42s)
* 05:04 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:03 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:01 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 05:01 tstarling@deploy1003: Scap cancelled without rolling back.
* 04:53 tstarling@deploy1003: tstarling: Continuing with deployment
* 04:29 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 04:25 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 43s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-01 ==
* 21:59 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] (duration: 18m 05s)
* 21:52 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:47 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]]
* 21:38 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] (duration: 23m 55s)
* 21:28 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:20 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:14 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]]
* 20:47 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1016.eqiad.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:24 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:11 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 20:11 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:57 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:57 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:56 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:54 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:52 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:47 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:45 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:42 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:41 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:40 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:39 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:32 jhancock@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudcephosd1055
* 19:32 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:30 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.* [reason: update IP addrs]
* 19:30 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 19:30 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:23 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:22 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:16 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:13 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1056
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1056
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1055
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:12 jclark@cumin1003: END (ERROR) - Cookbook sre.network.configure-switch-interfaces (exit_code=97) for host cloudcephosd1054
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1054
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:11 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:06 jclark@cumin1003: START - Cookbook sre.dns.netbox
* 19:06 sukhe@dns1004: END - running authdns-update
* 19:03 sukhe@dns1004: START - running authdns-update
* 18:14 dancy@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 17:04 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 17:00 dancy@deploy1003: Finished scap sync-world: testing (duration: 08m 07s)
* 16:52 dancy@deploy1003: Started scap sync-world: testing
* 16:48 dancy@deploy1003: sync-world aborted: testing (duration: 00m 05s)
* 16:48 dancy@deploy1003: Started scap sync-world: testing
* 16:47 dancy@deploy1003: Installation of scap version "4.287.0" completed for 156 hosts
* 16:42 dancy@deploy1003: Installing scap version "4.287.0" for 156 host(s)
* 16:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 16:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 15:51 moritzm: installing mesa security updates
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts phab1004.eqiad.wmnet
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:27 aokoth@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:20 aokoth@cumin1003: START - Cookbook sre.dns.netbox
* 15:14 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595] (duration: 05m 32s)
* 15:14 aokoth@cumin1003: START - Cookbook sre.hosts.decommission for hosts phab1004.eqiad.wmnet
* 15:11 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on phab1004.eqiad.wmnet with reason: Decom
* 15:09 joal@deploy1003: Started deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595]
* 15:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config: apply
* 15:08 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config: apply
* 14:55 hashar: Restarted Jenkins on releases1003
* 14:51 hashar: Restarted CI Jenkins on contint1003
* 14:48 hashar: Restarting Gerrit primary on gerrit2003
* 14:45 hashar: Restarted Gerrit on gerrit1003 and gerrit2002
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 14:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 14:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 14:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:24 moritzm: installing curl security updates
* 14:24 jmm@dns1004: END - running authdns-update
* 14:23 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser (duration: 00m 14s)
* 14:23 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:21 jmm@dns1004: START - running authdns-update
* 14:21 jmm@dns1004: END - running authdns-update
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 14:19 jmm@dns1004: START - running authdns-update
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet
* 14:14 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:14 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:13 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595] (duration: 07m 26s)
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Test
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Test
* 14:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet
* 14:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Test
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Test
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 14:05 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595]
* 14:05 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] (duration: 37m 37s)
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 hashar: Removed openjdk-17 packages from contint1002/contint2002 following relocation of CI Jenkins to contint1003/contint2003 # [[phab:T418521|T418521]]
* 14:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:02 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:02 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:00 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595] (duration: 00m 45s)
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 13:59 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595]
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:58 ladsgroup@dns1004: END - running authdns-update
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:57 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 ladsgroup@dns1004: START - running authdns-update
* 13:56 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:56 ladsgroup@dns1004: END - running authdns-update
* 13:55 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:53 ladsgroup@dns1004: START - running authdns-update
* 13:50 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:49 kharlan@deploy1003: kharlan: Continuing with deployment
* 13:48 kharlan@deploy1003: kharlan: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2004.wikimedia.org
* 13:42 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:41 jmm@dns1004: END - running authdns-update
* 13:40 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2004.wikimedia.org
* 13:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2003.wikimedia.org
* 13:38 jmm@dns1004: START - running authdns-update
* 13:34 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2003.wikimedia.org
* 13:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1004.wikimedia.org
* 13:26 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]]
* 13:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:25 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:24 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1004.wikimedia.org
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:23 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] (duration: 20m 24s)
* 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1003.wikimedia.org
* 13:20 fnegri@deploy1003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply
* 13:19 fnegri@deploy1003: helmfile [eqiad] START helmfile.d/services/toolhub: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:16 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1003.wikimedia.org
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:15 moritzm: bump urldownloader[12]00[34] to 8G RAM [[phab:T429175|T429175]]
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:15 fnegri@deploy1003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 fnegri@deploy1003: helmfile [codfw] START helmfile.d/services/toolhub: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:11 fnegri@deploy1003: helmfile [staging] DONE helmfile.d/services/toolhub: apply
* 13:11 aude@deploy1003: aude: Continuing with deployment
* 13:10 fnegri@deploy1003: helmfile [staging] START helmfile.d/services/toolhub: apply
* 13:07 aude@deploy1003: aude: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2004.wikimedia.org with OS bookworm
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow1004.eqiad.wmnet
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1004.eqiad.wmnet with OS trixie
* 12:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:47 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:41 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] (duration: 16m 25s)
* 12:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 12:34 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 12:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 12:33 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 12:31 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:29 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:25 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]]
* 12:22 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1004.eqiad.wmnet with OS trixie
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:20 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 12:20 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 12:16 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow1004.eqiad.wmnet
* 12:15 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2004.wikimedia.org with OS bookworm
* 12:14 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 12:07 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2003.wikimedia.org with OS bookworm
* 11:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:43 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:32 jmm@cumin2003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling restart_daemons on A:kafka-test-eqiad
* 11:26 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2003.wikimedia.org with OS bookworm
* 11:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1004.wikimedia.org with OS bookworm
* 11:12 moritzm: installing openjdk-21 security updates
* 11:12 jmm@cumin2003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling restart_daemons on A:kafka-test-eqiad
* 10:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:53 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:49 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2902: Pool back db2902
* 10:45 moritzm: installing Python 3.11 security updates
* 10:37 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1004.wikimedia.org with OS bookworm
* 10:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5022.eqsin.wmnet with OS trixie
* 10:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp5022.eqsin.wmnet on all recursors
* 10:36 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache cp5022.eqsin.wmnet on all recursors
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:30 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:04 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:04 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2902: Pool back db2902
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:03 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2902: test
* 10:03 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 10:01 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 09:58 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:50 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/echoserver: apply
* 09:49 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/echoserver: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Change db1176 and db2230's weight, test-s4 masters, to 0 to mimic the rest of production [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96292 and previous config saved to /var/cache/conftool/dbconfig/20260901-092444-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96291 and previous config saved to /var/cache/conftool/dbconfig/20260901-090233-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96290 and previous config saved to /var/cache/conftool/dbconfig/20260901-090158-marostegui.json
* 09:01 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96289 and previous config saved to /var/cache/conftool/dbconfig/20260901-090121-marostegui.json
* 09:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:56 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:50 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1003.wikimedia.org with OS bookworm
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1074.eqiad.wmnet
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:46 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:43 marostegui@cumin1003: dbctl commit (dc=all): 'Test repool db2902', diff saved to https://phabricator.wikimedia.org/P96288 and previous config saved to /var/cache/conftool/dbconfig/20260901-084317-marostegui.json
* 08:42 marostegui@cumin1003: dbctl commit (dc=all): 'Test depool db2902', diff saved to https://phabricator.wikimedia.org/P96287 and previous config saved to /var/cache/conftool/dbconfig/20260901-084249-marostegui.json
* 08:39 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:37 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2902: test
* 08:36 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96285 and previous config saved to /var/cache/conftool/dbconfig/20260901-083557-marostegui.json
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96284 and previous config saved to /var/cache/conftool/dbconfig/20260901-083527-marostegui.json
* 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96283 and previous config saved to /var/cache/conftool/dbconfig/20260901-083432-marostegui.json
* 08:32 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1074.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1073.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:27 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:24 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 08:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022']
* 08:16 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1073.eqiad.wmnet
* 08:15 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1003.wikimedia.org with OS bookworm
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1072.eqiad.wmnet
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:14 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:10 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:05 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1072.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1067.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:00 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:56 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:54 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022']
* 07:53 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 07:52 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie
* 07:52 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 07:50 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1067.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1066.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:47 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 07:41 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:34 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1066.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1065.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:32 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:27 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:23 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1065.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1075.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:15 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 06:49 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:45 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1075.eqiad.wmnet
* 06:29 moritzm: installing Java 17 security updates
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.15 (duration: 02m 25s)
* 03:50 denisse@deploy1003: Finished deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2 (duration: 00m 19s)
* 03:50 denisse@deploy1003: Started deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2
* 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]] (duration: 37m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:30 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 00:29 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
== Other archives ==
See [[Server Admin Log/Archives]].
<noinclude>
[[Category:SAL]]
[[Category:Operations]]
</noinclude>
nwy8xa664s64mvbto3bo3i5kqk1fwkh
2458703
2458702
2026-09-19T16:55:31Z
Stashbot
7414
ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
2458703
wikitext
text/x-wiki
== 2026-09-19 ==
* 16:55 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:11 urbanecm: Attach SHB@commonswiki to the SUL account manually ([[phab:T438591|T438591]], see [[phab:T438591|T438591]]#12341750 for what I did exactly)
* 04:08 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 04:08 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 04:08 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 04:07 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 36s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-18 ==
* 22:41 rzl@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=sessionstore,name=eqiad
* 17:08 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343065{{!}}mw.DesktopArticleTarget: if source education is enabled suppress welcome (T434249)]] (duration: 09m 26s)
* 17:05 Dreamy_Jazz: Created `securepoll_log` on `nlwiki` main DB cluster for [[phab:T434045|T434045]]
* 17:04 kemayo@deploy1003: kemayo: Continuing with deployment
* 17:03 kemayo@deploy1003: kemayo: Backport for [[gerrit:1343065{{!}}mw.DesktopArticleTarget: if source education is enabled suppress welcome (T434249)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:59 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1343065{{!}}mw.DesktopArticleTarget: if source education is enabled suppress welcome (T434249)]]
* 16:49 oblivian@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343127{{!}}ResourceLoader: hotfix for current logspam over the weekend (T438387)]] (duration: 11m 36s)
* 16:42 oblivian@deploy1003: oblivian: Continuing with deployment
* 16:42 oblivian@deploy1003: oblivian: Backport for [[gerrit:1343127{{!}}ResourceLoader: hotfix for current logspam over the weekend (T438387)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:37 oblivian@deploy1003: Started scap sync-world: Backport for [[gerrit:1343127{{!}}ResourceLoader: hotfix for current logspam over the weekend (T438387)]]
* 16:09 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ml-serve1016.eqiad.wmnet with OS trixie
* 14:49 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie
* 13:37 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host registry1004.eqiad.wmnet with OS trixie
* 13:23 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on registry1004.eqiad.wmnet with reason: host reimage
* 13:18 elukey@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on registry1004.eqiad.wmnet with reason: host reimage
* 13:04 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host registry1004.eqiad.wmnet with OS trixie
* 12:14 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:13 jclark@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:30 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 09:30 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 09:22 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on an-redacteddb1001.eqiad.wmnet with reason: cloning
* 09:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 09:19 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 09:18 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 21 hosts with reason: cloning db1270
* 09:18 marostegui: clone db1270:x4 from db1155:x4 lag will appear on x4
* 09:09 brouberol@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 09:08 brouberol@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 08:23 brouberol@dns1004: END - running authdns-update
* 08:21 brouberol@dns1004: START - running authdns-update
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 57s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-17 ==
* 21:04 tsev@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T438395|T438395]]
* 20:55 tsev@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T438395|T438395]]
* 20:47 jhuneidi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342774{{!}}Worklist Promotion test kitchen - Enable flag in production (T434513)]], [[gerrit:1342781{{!}}Exclude returntoapp query from app interception on iOS (T438395)]], [[gerrit:1342798{{!}}Revert "Update wikimania wordmark for 2026"]] (duration: 35m 59s)
* 20:35 jhuneidi@deploy1003: robertsky, jhuneidi, cmelo, tsev: Continuing with deployment
* 20:31 jhuneidi@deploy1003: robertsky, jhuneidi, cmelo, tsev: Backport for [[gerrit:1342774{{!}}Worklist Promotion test kitchen - Enable flag in production (T434513)]], [[gerrit:1342781{{!}}Exclude returntoapp query from app interception on iOS (T438395)]], [[gerrit:1342798{{!}}Revert "Update wikimania wordmark for 2026"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:11 jhuneidi@deploy1003: Started scap sync-world: Backport for [[gerrit:1342774{{!}}Worklist Promotion test kitchen - Enable flag in production (T434513)]], [[gerrit:1342781{{!}}Exclude returntoapp query from app interception on iOS (T438395)]], [[gerrit:1342798{{!}}Revert "Update wikimania wordmark for 2026"]]
* 19:15 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1100.eqiad.wmnet with OS trixie
* 19:15 vriley@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1004"
* 19:10 vriley@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1004"
* 18:52 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1100.eqiad.wmnet with reason: host reimage
* 18:48 vriley@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1100.eqiad.wmnet with reason: host reimage
* 18:32 vriley@cumin1004: START - Cookbook sre.hosts.reimage for host ms-be1100.eqiad.wmnet with OS trixie
* 18:20 urbanecm: Deploy a security fix for [[phab:T438389|T438389]]
* 17:47 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be1100.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:38 vriley@cumin1004: START - Cookbook sre.hosts.provision for host ms-be1100.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:38 vriley@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1100.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:37 vriley@cumin1004: START - Cookbook sre.hosts.provision for host ms-be1100.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:23 vriley@cumin1004: START - Cookbook sre.hosts.reimage for host ms-be1100.eqiad.wmnet with OS trixie
* 16:52 aokoth@deploy1003: Finished deploy [phabricator/deployment@c386249]: Deploy Phab (duration: 00m 12s)
* 16:52 aokoth@deploy1003: Started deploy [phabricator/deployment@c386249]: Deploy Phab
* 16:42 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1099.eqiad.wmnet with OS trixie
* 16:42 vriley@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1004"
* 16:42 vriley@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1004"
* 16:35 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be1100.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:31 aokoth@deploy1003: Finished deploy [phabricator/deployment@c386249]: Deploy Phab (duration: 00m 19s)
* 16:31 aokoth@deploy1003: Started deploy [phabricator/deployment@c386249]: Deploy Phab
* 16:21 vriley@cumin1004: START - Cookbook sre.hosts.provision for host ms-be1100.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:20 vriley@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be1100
* 16:20 vriley@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host ms-be1100
* 16:19 vriley@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:19 vriley@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [ms-be1100] - vriley@cumin1004"
* 16:19 vriley@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [ms-be1100] - vriley@cumin1004"
* 16:15 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 16:15 vriley@cumin1004: START - Cookbook sre.dns.netbox
* 16:15 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 16:14 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 16:14 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 16:12 moritzm: installing libapache-mod-auth-oidc security updates
* 16:12 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1099.eqiad.wmnet with reason: host reimage
* 16:11 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 16:11 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 16:08 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/editcheck-headless: apply
* 16:07 rzl@deploy1003: helmfile [staging] START helmfile.d/services/editcheck-headless: apply
* 16:06 vriley@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1099.eqiad.wmnet with reason: host reimage
* 16:01 moritzm: installing aom security updates
* 16:01 btullis@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ceph-admin2001.codfw.wmnet
* 16:01 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ceph-admin2001.codfw.wmnet with OS bookworm
* 15:55 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 15:55 rzl@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 15:54 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 15:54 rzl@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 15:54 rzl@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 15:53 rzl@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 15:53 rzl@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 15:52 rzl@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 15:51 vriley@cumin1004: START - Cookbook sre.hosts.reimage for host ms-be1099.eqiad.wmnet with OS trixie
* 15:48 moritzm: installing libde265 security updates
* 15:44 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ceph-admin2001.codfw.wmnet with reason: host reimage
* 15:40 btullis@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ceph-admin1001.eqiad.wmnet
* 15:40 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ceph-admin1001.eqiad.wmnet with OS bookworm
* 15:39 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir4003.*
* 15:37 btullis@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ceph-admin2001.codfw.wmnet with reason: host reimage
* 15:26 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir4003.ulsfo.wmnet with OS trixie
* 15:25 aokoth@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host phab2003.codfw.wmnet with OS trixie
* 15:23 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ceph-admin1001.eqiad.wmnet with reason: host reimage
* 15:19 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host registry1005.eqiad.wmnet with OS trixie
* 15:17 btullis@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ceph-admin1001.eqiad.wmnet with reason: host reimage
* 15:16 btullis@cumin1004: START - Cookbook sre.hosts.reimage for host ceph-admin2001.codfw.wmnet with OS bookworm
* 15:16 btullis@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ceph-admin2001.codfw.wmnet - btullis@cumin1004"
* 15:16 btullis@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ceph-admin2001.codfw.wmnet - btullis@cumin1004"
* 15:15 btullis@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ceph-admin2001.codfw.wmnet on all recursors
* 15:15 btullis@cumin1004: START - Cookbook sre.dns.wipe-cache ceph-admin2001.codfw.wmnet on all recursors
* 15:15 btullis@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:15 btullis@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ceph-admin2001.codfw.wmnet - btullis@cumin1004"
* 15:15 btullis@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ceph-admin2001.codfw.wmnet - btullis@cumin1004"
* 15:08 aokoth@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on phab2003.codfw.wmnet with reason: host reimage
* 15:06 Msz2001: Deployed private code changes to Suggestedinvestigations
* 15:05 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir4003.ulsfo.wmnet with reason: host reimage
* 15:05 aokoth@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on phab2003.codfw.wmnet with reason: host reimage
* 15:04 btullis@cumin1004: START - Cookbook sre.hosts.reimage for host ceph-admin1001.eqiad.wmnet with OS bookworm
* 15:04 btullis@cumin1004: START - Cookbook sre.dns.netbox
* 15:04 btullis@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ceph-admin1001.eqiad.wmnet - btullis@cumin1004"
* 15:04 btullis@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ceph-admin1001.eqiad.wmnet - btullis@cumin1004"
* 15:04 btullis@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ceph-admin1001.eqiad.wmnet on all recursors
* 15:04 btullis@cumin1004: START - Cookbook sre.dns.wipe-cache ceph-admin1001.eqiad.wmnet on all recursors
* 15:04 btullis@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:04 btullis@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ceph-admin1001.eqiad.wmnet - btullis@cumin1004"
* 15:04 btullis@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ceph-admin1001.eqiad.wmnet - btullis@cumin1004"
* 15:02 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on registry1005.eqiad.wmnet with reason: host reimage
* 15:01 btullis@cumin1004: START - Cookbook sre.ganeti.makevm for new host ceph-admin2001.codfw.wmnet
* 15:00 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342265{{!}}JsonSchemaBuilder: Cache the root schema in the process (T437588)]], [[gerrit:1342264{{!}}JsonSchemaBuilder: Cache the root schema in the process (T437588)]], [[gerrit:1342684{{!}}SI: Preserve the username filter when switching queues (T438308)]] (duration: 12m 34s)
* 15:00 btullis@cumin1004: START - Cookbook sre.dns.netbox
* 15:00 btullis@cumin1004: START - Cookbook sre.ganeti.makevm for new host ceph-admin1001.eqiad.wmnet
* 14:58 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir4003.ulsfo.wmnet with reason: host reimage
* 14:58 elukey@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on registry1005.eqiad.wmnet with reason: host reimage
* 14:55 urbanecm@deploy1003: mszwarc, urbanecm: Continuing with deployment
* 14:51 urbanecm@deploy1003: mszwarc, urbanecm: Backport for [[gerrit:1342265{{!}}JsonSchemaBuilder: Cache the root schema in the process (T437588)]], [[gerrit:1342264{{!}}JsonSchemaBuilder: Cache the root schema in the process (T437588)]], [[gerrit:1342684{{!}}SI: Preserve the username filter when switching queues (T438308)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:51 aokoth@cumin1004: START - Cookbook sre.hosts.reimage for host phab2003.codfw.wmnet with OS trixie
* 14:50 aokoth@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on phab2003.codfw.wmnet with reason: Reimage
* 14:47 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1342265{{!}}JsonSchemaBuilder: Cache the root schema in the process (T437588)]], [[gerrit:1342264{{!}}JsonSchemaBuilder: Cache the root schema in the process (T437588)]], [[gerrit:1342684{{!}}SI: Preserve the username filter when switching queues (T438308)]]
* 14:42 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1331830{{!}}Keep Balinese Palm Leaf variants enabled on wikisource (T436398)]], [[gerrit:1340216{{!}}Turn on variant conversion for PageAssessments (T328012)]], [[gerrit:1341949{{!}}Parsoid Read Views: Enable on 61 wikiquote wikis (T437917)]] (duration: 15m 31s)
* 14:39 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host registry1005.eqiad.wmnet with OS trixie
* 14:35 cscott@deploy1003: ssastry, cscott: Continuing with deployment
* 14:33 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir4003.ulsfo.wmnet with OS trixie
* 14:32 cscott@deploy1003: ssastry, cscott: Backport for [[gerrit:1331830{{!}}Keep Balinese Palm Leaf variants enabled on wikisource (T436398)]], [[gerrit:1340216{{!}}Turn on variant conversion for PageAssessments (T328012)]], [[gerrit:1341949{{!}}Parsoid Read Views: Enable on 61 wikiquote wikis (T437917)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:26 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1331830{{!}}Keep Balinese Palm Leaf variants enabled on wikisource (T436398)]], [[gerrit:1340216{{!}}Turn on variant conversion for PageAssessments (T328012)]], [[gerrit:1341949{{!}}Parsoid Read Views: Enable on 61 wikiquote wikis (T437917)]]
* 14:20 caro@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342678{{!}}enwiki desktop VE: add education popup for switching to source editor (T434249)]], [[gerrit:1342362{{!}}Make VE the default editor on enwiki desktop (T436574)]] (duration: 33m 52s)
* 14:07 caro@deploy1003: caro: Continuing with deployment
* 14:06 caro@deploy1003: caro: Backport for [[gerrit:1342678{{!}}enwiki desktop VE: add education popup for switching to source editor (T434249)]], [[gerrit:1342362{{!}}Make VE the default editor on enwiki desktop (T436574)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:46 caro@deploy1003: Started scap sync-world: Backport for [[gerrit:1342678{{!}}enwiki desktop VE: add education popup for switching to source editor (T434249)]], [[gerrit:1342362{{!}}Make VE the default editor on enwiki desktop (T436574)]]
* 13:37 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 13:34 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:33 elukey@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:31 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:30 elukey@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:29 Emperor: apus - radosgw-admin quota set --quota-scope=user --uid=docker-registry --max-size=5T [[phab:T438339|T438339]]
* 13:24 jclark@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1016
* 13:24 jclark@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1016
* 13:11 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:11 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:09 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:09 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:55 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:53 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 12:42 jclark@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:40 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:40 jclark@cumin1003: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:35 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 12:19 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342626{{!}}MathMathML: Simplify Mathoid fallback/a11y class logic (T436026)]], [[gerrit:1342627{{!}}ext.math.mathjax: Implement mwe-math-mathml-a11y for client-side MathJax (T436026)]] (duration: 13m 04s)
* 12:14 krinkle@deploy1003: krinkle: Continuing with deployment
* 12:10 krinkle@deploy1003: krinkle: Backport for [[gerrit:1342626{{!}}MathMathML: Simplify Mathoid fallback/a11y class logic (T436026)]], [[gerrit:1342627{{!}}ext.math.mathjax: Implement mwe-math-mathml-a11y for client-side MathJax (T436026)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:05 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1342626{{!}}MathMathML: Simplify Mathoid fallback/a11y class logic (T436026)]], [[gerrit:1342627{{!}}ext.math.mathjax: Implement mwe-math-mathml-a11y for client-side MathJax (T436026)]]
* 10:37 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 10:28 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 10:10 blake@deploy1003: Finished scap sync-world: cleanup for [[phab:T417800|T417800]] (duration: 03m 57s)
* 10:07 blake@deploy1003: Started scap sync-world: cleanup for [[phab:T417800|T417800]]
* 09:52 marostegui@cumin1004: dbctl commit (dc=all): 'Fix weights [[phab:T436496|T436496]]', diff saved to https://phabricator.wikimedia.org/P96466 and previous config saved to /var/cache/conftool/dbconfig/20260917-095235-marostegui.json
* 09:51 marostegui@cumin1004: dbctl commit (dc=all): 'Fix weights [[phab:T436496|T436496]]', diff saved to https://phabricator.wikimedia.org/P96465 and previous config saved to /var/cache/conftool/dbconfig/20260917-095131-marostegui.json
* 09:41 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:41 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 09:41 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:40 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 09:40 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:40 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 09:35 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 09:35 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 08:37 moritzm: pruned obsolete Bullseye image prometheus-nutcracker-exporter from the docker registry [[phab:T416452|T416452]]
* 08:34 XioNoX: Manually install gnmic 0.49.0 on netflow2005 - [[phab:T438291|T438291]]
* 08:28 brouberol@dns1004: END - running authdns-update
* 08:26 brouberol@dns1004: START - running authdns-update
* 08:13 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.20 refs [[phab:T430839|T430839]]
* 08:10 moritzm: imported nodejs_26.8.2-1nodesource1 to thirdparty/node26 for trixie-wikimedia [[phab:T437510|T437510]]
* 08:07 Amir1: dropped links tables from db2206 ([[phab:T437278|T437278]])
* 08:03 Amir1: dropped links tables from db2219 ([[phab:T437278|T437278]])
* 08:01 Amir1: dropped links tables from db2236 ([[phab:T437278|T437278]])
* 07:59 Amir1: dropped non-links tables from db1262 ([[phab:T437278|T437278]])
* 07:57 Amir1: dropped non-links tables from db2245 ([[phab:T437278|T437278]])
* 07:52 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@8eaca67] (releasing): [[phab:T438205|T438205]] to prod host (duration: 00m 44s)
* 07:52 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@8eaca67] (releasing): [[phab:T438205|T438205]] to prod host
* 07:49 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@8eaca67] (releasing): [[phab:T438205|T438205]] to backup host (duration: 00m 47s)
* 07:48 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@8eaca67] (releasing): [[phab:T438205|T438205]] to backup host
* 07:25 XioNoX: Manually install gnmic 0.49.0 on netflow1004 - [[phab:T438291|T438291]]
* 07:24 mlitn@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342398{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342396{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342395{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342545{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342546{{!}}Localisation updates from https://translatewiki.net.]],
* 07:19 mlitn@deploy1003: mlitn, jdlrobson: Continuing with deployment
* {{safesubst:SAL entry|1=07:18 mlitn@deploy1003: mlitn, jdlrobson: Backport for [[gerrit:1342398{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342396{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342395{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342545{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342546{{!}}Localisation updates from https://translatewiki.net.]], [[gerri}}
* 07:11 mlitn@deploy1003: Started scap sync-world: Backport for [[gerrit:1342398{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342396{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342395{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342545{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342546{{!}}Localisation updates from https://translatewiki.net.]],
* 07:10 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Jmoore111 out of all services on: 2444 hosts
* 06:07 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 05:53 marostegui@cumin1004: END (FAIL) - Cookbook sre.mysql.decommission (exit_code=99)
* 05:53 marostegui@cumin1004: Removing db1180 from zarcillo [[phab:T437222|T437222]]
* 05:53 marostegui@cumin1004: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1180.eqiad.wmnet
* 05:53 marostegui@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 05:53 marostegui@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1180.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1004"
* 05:53 marostegui@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1180.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1004"
* 05:49 marostegui@cumin1004: START - Cookbook sre.dns.netbox
* 05:44 marostegui@cumin1004: START - Cookbook sre.hosts.decommission for hosts db1180.eqiad.wmnet
* 05:43 marostegui@cumin1004: START - Cookbook sre.mysql.decommission
* 04:26 aokoth@cumin1004: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet
* 04:24 aokoth@cumin1004: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet
== 2026-09-16 ==
* 23:10 rzl: rzl@deploy1003 Finished scap sync-world: Backport for [[gerrit:1342091{{!}}Repool poolcounter[1007,2006] (T435163)]] (duration: 11m 09s)
* 22:50 rzl@deploy1003: Started scap sync-world: Backport for [[gerrit:1342091{{!}}Repool poolcounter[1007,2006] (T435163)]]
* 22:47 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342383{{!}}DonorIdentification: Confirm before unlinking donor status in preferences (T436698)]], [[gerrit:1342385{{!}}Make learn more link to new window (T438252)]] (duration: 35m 21s)
* 22:35 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 22:33 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1342383{{!}}DonorIdentification: Confirm before unlinking donor status in preferences (T436698)]], [[gerrit:1342385{{!}}Make learn more link to new window (T438252)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:12 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1342383{{!}}DonorIdentification: Confirm before unlinking donor status in preferences (T436698)]], [[gerrit:1342385{{!}}Make learn more link to new window (T438252)]]
* 22:10 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2006.codfw.wmnet
* 22:06 rzl@cumin2003: START - Cookbook sre.hosts.reboot-single for host poolcounter2006.codfw.wmnet
* 22:06 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1007.eqiad.wmnet
* 22:02 rzl@cumin2003: START - Cookbook sre.hosts.reboot-single for host poolcounter1007.eqiad.wmnet
* 21:56 rzl@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342090{{!}}Repool poolcounter[1006,2005]; depool poolcounter[1007,2006] for reboot (T435163)]] (duration: 09m 39s)
* 21:52 rzl@deploy1003: rzl: Continuing with deployment
* 21:51 rzl@deploy1003: rzl: Backport for [[gerrit:1342090{{!}}Repool poolcounter[1006,2005]; depool poolcounter[1007,2006] for reboot (T435163)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:47 rzl@deploy1003: Started scap sync-world: Backport for [[gerrit:1342090{{!}}Repool poolcounter[1006,2005]; depool poolcounter[1007,2006] for reboot (T435163)]]
* 21:46 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:43 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2005.codfw.wmnet
* 21:42 vriley@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1099.eqiad.wmnet with OS trixie
* 21:41 vriley@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1098.eqiad.wmnet with OS trixie
* 21:41 vriley@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin2003"
* 21:40 vriley@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin2003"
* 21:40 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:40 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:39 rzl@cumin2003: START - Cookbook sre.hosts.reboot-single for host poolcounter2005.codfw.wmnet
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1006.eqiad.wmnet
* 21:38 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:38 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:37 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:35 rzl@cumin2003: START - Cookbook sre.hosts.reboot-single for host poolcounter1006.eqiad.wmnet
* 21:32 rzl@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342089{{!}}Depool poolcounter[1006,2005] for reboot (T435163)]] (duration: 13m 53s)
* 21:26 rzl@deploy1003: rzl: Continuing with deployment
* 21:25 rzl@deploy1003: rzl: Backport for [[gerrit:1342089{{!}}Depool poolcounter[1006,2005] for reboot (T435163)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:23 vriley@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1098.eqiad.wmnet with reason: host reimage
* 21:18 rzl@deploy1003: Started scap sync-world: Backport for [[gerrit:1342089{{!}}Depool poolcounter[1006,2005] for reboot (T435163)]]
* 21:17 vriley@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1098.eqiad.wmnet with reason: host reimage
* 21:10 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:09 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:09 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:09 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:08 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:02 vriley@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS trixie
* 20:49 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342339{{!}}Reapply "Tell VisualEditor about the app web edit tags", modified]] (duration: 35m 55s)
* 20:37 kemayo@deploy1003: cklimas, kemayo: Continuing with deployment
* 20:33 kemayo@deploy1003: cklimas, kemayo: Backport for [[gerrit:1342339{{!}}Reapply "Tell VisualEditor about the app web edit tags", modified]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:13 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1342339{{!}}Reapply "Tell VisualEditor about the app web edit tags", modified]]
* 19:20 dwisehaupt@dns1005: END - running authdns-update
* 19:18 dwisehaupt@dns1005: START - running authdns-update
* 19:06 vriley@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:57 dwisehaupt@dns1005: END - running authdns-update
* 18:55 vriley@cumin2003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:55 dwisehaupt@dns1005: START - running authdns-update
* 18:44 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 18:42 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 18:37 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 18:35 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 18:26 robh@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:22 robh@cumin2003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:21 dzahn@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 18:21 robh@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:21 robh@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:20 dzahn@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 18:20 dzahn@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 18:18 dzahn@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 18:18 mutante: k8s/miscweb: admin_ng deploy: creating namespace for attribution.wikimedia.org [[phab:T437635|T437635]]
* 18:17 dzahn@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 18:17 dzahn@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 18:17 dzahn@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 18:16 dzahn@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 18:13 cdanis@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "fix known-client creation - cdanis@cumin1003"
* 18:13 cdanis@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: fix known-client creation - cdanis@cumin1003
* 18:12 cdanis@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: fix known-client creation - cdanis@cumin1003
* 18:12 cdanis@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "fix known-client creation - cdanis@cumin1003"
* 18:04 vriley@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:51 vriley@cumin2003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:46 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 17:46 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 17:45 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:45 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:44 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:44 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:35 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1006.eqiad.wmnet with OS trixie
* 17:35 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1003"
* 17:30 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:29 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1003"
* 17:27 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:27 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:23 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:23 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:20 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 17:20 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 17:14 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:13 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1006.eqiad.wmnet with reason: host reimage
* 17:10 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:10 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:10 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:10 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:09 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:07 jclark@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1006.eqiad.wmnet with reason: host reimage
* 17:07 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 17:05 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:05 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:52 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie
* 16:46 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir7003.magru.wmnet
* 16:44 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir7003
* 16:19 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie
* 15:42 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply
* 15:42 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply
* 15:42 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 15:42 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir7003.magru.wmnet with OS trixie
* 15:41 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 15:36 jnuche@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342279{{!}}Use parser output value instead of status (T438154)]] (duration: 33m 21s)
* 15:35 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply
* 15:35 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply
* 15:24 jnuche@deploy1003: jnuche, jforrester: Continuing with deployment
* 15:23 jnuche@deploy1003: jnuche, jforrester: Backport for [[gerrit:1342279{{!}}Use parser output value instead of status (T438154)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:19 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie
* 15:18 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir7003.magru.wmnet with reason: host reimage
* 15:14 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir7003.magru.wmnet with reason: host reimage
* 15:03 jnuche@deploy1003: Started scap sync-world: Backport for [[gerrit:1342279{{!}}Use parser output value instead of status (T438154)]]
* 14:50 moritzm: installing apache2 security updates
* 14:49 slyngshede@cumin1003: conftool action : set/pooled=yes; selector: name=cp5026.eqsin.wmnet
* 14:47 slyngshede@cumin1003: conftool action : set/weight=1; selector: name=cp5026.eqsin.wmnet
* 14:45 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5026.eqsin.wmnet with OS trixie
* 14:44 moritzm: installing python-filelock security updates
* 14:42 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir7003.magru.wmnet with OS trixie
* 14:35 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:35 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:34 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:33 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp6002.drmrs.wmnet
* 14:32 sukhe@puppetserver1001: conftool action : set/weight=100; selector: name=cp6002.drmrs.wmnet,service=ats-be
* 14:32 sukhe@puppetserver1001: conftool action : set/weight=1; selector: name=cp6002.drmrs.wmnet,service=cdn
* 14:29 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6002.drmrs.wmnet with OS trixie
* 14:24 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: sync
* 14:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:24 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:24 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: sync
* 14:14 jforrester@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.19,1.47.0-wmf.20,next --multiversion-image-basename docker-registry.discovery.wmnet/restricte
* 14:14 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:14 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:13 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1342008{{!}}abstractwiki: Add three new articles per community advice to show off the feature (T434227)]]
* 14:10 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: sync
* 14:10 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: sync
* 14:10 jforrester@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.19,1.47.0-wmf.20,next --multiversion-image-basename docker-registry.discovery.wmnet/restricte
* 14:10 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: sync
* 14:10 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: sync
* 14:09 Amir1: dropped links tables on db2237 ([[phab:T437278|T437278]])
* 14:08 Amir1: dropped links tables on db1238 ([[phab:T437278|T437278]])
* 14:07 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1342008{{!}}abstractwiki: Add three new articles per community advice to show off the feature (T434227)]]
* 14:03 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5026.eqsin.wmnet with reason: host reimage
* 14:02 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:02 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:02 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 14:01 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:00 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 13:59 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6002.drmrs.wmnet with reason: host reimage
* 13:56 slyngshede@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5026.eqsin.wmnet with reason: host reimage
* 13:54 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cp6002.drmrs.wmnet with reason: host reimage
* 13:53 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 13:52 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 13:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dse-k8s-etcd[1001-1003].eqiad.wmnet
* 13:48 bking@cumin2003: START - Cookbook sre.hosts.remove-downtime for dse-k8s-etcd[1001-1003].eqiad.wmnet
* 13:46 bking@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM dse-k8s-etcd1001.eqiad.wmnet
* 13:46 bking@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM dse-k8s-etcd1001.eqiad.wmnet
* 13:45 bking@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM dse-k8s-etcd1002.eqiad.wmnet
* 13:41 bking@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM dse-k8s-etcd1002.eqiad.wmnet
* 13:41 bking@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM dse-k8s-etcd1003.eqiad.wmnet
* 13:38 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp6002.drmrs.wmnet with OS trixie
* 13:37 bking@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM dse-k8s-etcd1003.eqiad.wmnet
* 13:37 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.reboot-vm (exit_code=99) for VM dse-k8s-etcd1003.eqiad.wmnet
* 13:37 bking@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM dse-k8s-etcd1003.eqiad.wmnet
* 13:34 slyngshede@cumin1003: START - Cookbook sre.hosts.reimage for host cp5026.eqsin.wmnet with OS trixie
* 13:34 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 13:33 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5026.mgmt.eqsin.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:29 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp6002.mgmt.drmrs.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:25 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on dse-k8s-etcd[1001-1003].eqiad.wmnet with reason: Maintenance to increase vCPUS [[phab:T438084|T438084]]
* 13:24 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:24 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:22 slyngshede@cumin1003: START - Cookbook sre.hosts.provision for host cp5026.mgmt.eqsin.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342231{{!}}SI: Unset all filters on links to cases (T434530)]], [[gerrit:1342234{{!}}SI: Unset all filters on links to cases (T434530)]] (duration: 13m 10s)
* 13:19 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org
* 13:19 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org
* 13:19 sukhe@cumin1004: START - Cookbook sre.hosts.provision for host cp6002.mgmt.drmrs.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:18 stran@deploy1003: stran: Continuing with deployment
* 13:15 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: apply
* 13:15 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: apply
* 13:13 stran@deploy1003: stran: Backport for [[gerrit:1342231{{!}}SI: Unset all filters on links to cases (T434530)]], [[gerrit:1342234{{!}}SI: Unset all filters on links to cases (T434530)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:11 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp5026.eqsin.wmnet
* 13:09 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1342231{{!}}SI: Unset all filters on links to cases (T434530)]], [[gerrit:1342234{{!}}SI: Unset all filters on links to cases (T434530)]]
* 13:09 sukhe@cumin1004: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp6002.drmrs.wmnet
* 13:05 sukhe@cumin1004: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp6002.drmrs.wmnet
* 13:05 sukhe@cumin1004: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts cp6002.drmrs.wmnet
* 13:00 dkertesz@cumin1004: conftool action : set/pooled=yes; selector: name=cp5025.eqsin.wmnet
* 12:59 dkertesz@cumin1004: conftool action : set/weight=1; selector: name=cp5025.eqsin.wmnet
* 12:52 sukhe@cumin1004: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp6002.drmrs.wmnet
* 12:52 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cp6002.mgmt.drmrs.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:51 dkertesz: eqsin pooled again ([[phab:T438052|T438052]])
* 12:49 dkertesz@cumin1004: conftool action : set/pooled=yes; selector: cluster=dnsbox,dc=eqsin
* 12:47 dkertesz@dns1004: END - running authdns-update
* 12:45 dkertesz@dns1004: START - running authdns-update
* 12:43 dkertesz@cumin1004: conftool action : set/pooled=yes; selector: cluster=dnsbox,dc=eqsin,service=authdns-update
* 12:41 dkertesz@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqsin [reason: no reason specified, [[phab:T438052|T438052]]]
* 12:41 dkertesz@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool eqsin [reason: no reason specified, [[phab:T438052|T438052]]]
* 12:38 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a1-eqiad
* 12:38 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a1-eqiad
* 12:34 sukhe@cumin1004: START - Cookbook sre.hosts.provision for host cp6002.mgmt.drmrs.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:34 sukhe@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cp6002.drmrs.wmnet with reason: reimage
* 12:33 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp6002.drmrs.wmnet
* 12:13 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5025.eqsin.wmnet with OS trixie
* 12:12 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 12:11 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-a1-eqiad
* 12:09 cmooney@cumin1004: START - Cookbook sre.network.tls for network device ssw1-a1-eqiad
* 12:01 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.20 refs [[phab:T430839|T430839]]
* 11:59 moritzm: pruned obsolete Bullseye image python3-bullseye from the docker registry [[phab:T416452|T416452]]
* 11:50 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341284{{!}}IS/IS-labs: Set wmgUseModeratorToolkit default false (T431000)]] (duration: 10m 32s)
* 11:46 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ml-lab1002.eqiad.wmnet
* 11:45 samtar@deploy1003: samtar: Continuing with deployment
* 11:44 samtar@deploy1003: samtar: Backport for [[gerrit:1341284{{!}}IS/IS-labs: Set wmgUseModeratorToolkit default false (T431000)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:41 klausman@cumin1003: START - Cookbook sre.hosts.reboot-single for host ml-lab1002.eqiad.wmnet
* 11:39 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1341284{{!}}IS/IS-labs: Set wmgUseModeratorToolkit default false (T431000)]]
* 11:39 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5025.eqsin.wmnet with reason: host reimage
* 11:35 slyngshede@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5025.eqsin.wmnet with reason: host reimage
* 11:34 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:33 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:31 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:31 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:27 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:27 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:26 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:25 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:24 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.20 refs [[phab:T430839|T430839]]
* 11:21 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply
* 11:21 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply
* 11:20 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply
* 11:20 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply
* 11:19 moritzm: kicked off a new run of production-images-weekly-rebuild.service on build2004 (previously some leftovers of buster in the config prevented a complete run)
* 11:17 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
* 11:16 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
* 11:11 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
* 11:11 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
* 11:10 jnuche@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342210{{!}}Revert "Tell VisualEditor about the app web edit tags" (T437736 T438125)]] (duration: 33m 14s)
* 11:10 slyngshede@cumin1003: START - Cookbook sre.hosts.reimage for host cp5025.eqsin.wmnet with OS trixie
* 11:05 marostegui@cumin1004: dbctl commit (dc=all): 'Remove db1180 from dbctl [[phab:T437222|T437222]]', diff saved to https://phabricator.wikimedia.org/P96459 and previous config saved to /var/cache/conftool/dbconfig/20260916-110502-marostegui.json
* 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 11:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:01 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 11:01 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 10:57 jnuche@deploy1003: jnuche: Continuing with deployment
* 10:57 jnuche@deploy1003: jnuche: Backport for [[gerrit:1342210{{!}}Revert "Tell VisualEditor about the app web edit tags" (T437736 T438125)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:37 jnuche@deploy1003: Started scap sync-world: Backport for [[gerrit:1342210{{!}}Revert "Tell VisualEditor about the app web edit tags" (T437736 T438125)]]
* 10:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5025.mgmt.eqsin.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:10 slyngshede@cumin1003: START - Cookbook sre.hosts.provision for host cp5025.mgmt.eqsin.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:02 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 10:02 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 09:58 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 09:49 slyngshede@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on cp5025.eqsin.wmnet with reason: reimaging
* 09:48 slyngshede@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cp5025.eqsin.wmnet with reason: reimaging
* 09:41 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-timeline: apply
* 09:41 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-timeline: apply
* 09:38 moritzm: imported routinator 0.15.2-1trixie to thirdparty/routinator [[phab:T438122|T438122]]
* 09:30 slyngshede@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cp5025.mgmt.eqsin.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:29 slyngshede@cumin1003: START - Cookbook sre.hosts.provision for host cp5025.mgmt.eqsin.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:19 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 09:19 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp5025.eqsin.wmnet
* 09:19 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'sync'.
* 09:18 slyngshede@cumin1003: conftool action : set/pooled=yes; selector: name=cp3074.esams.wmnet
* 09:18 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp3074.esams.wmnet
* 09:15 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 09:12 elukey@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'sync'.
* 09:12 elukey@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'sync'.
* 09:11 elukey@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'sync'.
* 09:11 elukey@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'sync'.
* 09:10 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 09:10 elukey@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 09:09 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 09:09 elukey@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 08:55 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 08:54 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 08:40 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:40 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:36 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqsin [reason: depooling for maintainance, [[phab:T438052|T438052]]]
* 08:36 slyngshede@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool eqsin [reason: depooling for maintainance, [[phab:T438052|T438052]]]
* 08:35 slyngshede@cumin1003: END (FAIL) - Cookbook sre.dns.admin (exit_code=99) DNS admin: depool eqsin [reason: no reason specified, no task ID specified]
* 08:35 slyngshede@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool eqsin [reason: no reason specified, no task ID specified]
* 08:35 slyngshede@cumin1003: conftool action : set/pooled=no; selector: cluster=dnsbox,dc=eqsin
* 08:34 fabfur: start depooling eqsin ([[phab:T438052|T438052]])
* 08:24 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.20 refs [[phab:T430839|T430839]]
* 08:22 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:22 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:14 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.20 refs [[phab:T430839|T430839]]
* 08:11 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 08:11 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 08:11 jnuche@deploy1003: Rolling back deployment
* 08:10 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 08:07 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 07:59 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 07:59 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 07:58 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 07:54 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 07:34 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply
* 07:34 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply
* 07:20 mlitn@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342112{{!}}Adds an instrument for pre-image-carousel-retest (T437076)]], [[gerrit:1342113{{!}}Adds an instrument for pre-image-carousel-retest (T437076)]], [[gerrit:1342117{{!}}Set up instrument for 5-arm test (T437076)]], [[gerrit:1342118{{!}}Set up instrument for 5-arm test (T437076)]] (duration: 10m 56s)
* 07:16 mlitn@deploy1003: mlitn: Continuing with deployment
* 07:15 mlitn@deploy1003: mlitn: Backport for [[gerrit:1342112{{!}}Adds an instrument for pre-image-carousel-retest (T437076)]], [[gerrit:1342113{{!}}Adds an instrument for pre-image-carousel-retest (T437076)]], [[gerrit:1342117{{!}}Set up instrument for 5-arm test (T437076)]], [[gerrit:1342118{{!}}Set up instrument for 5-arm test (T437076)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be veri
* 07:09 mlitn@deploy1003: Started scap sync-world: Backport for [[gerrit:1342112{{!}}Adds an instrument for pre-image-carousel-retest (T437076)]], [[gerrit:1342113{{!}}Adds an instrument for pre-image-carousel-retest (T437076)]], [[gerrit:1342117{{!}}Set up instrument for 5-arm test (T437076)]], [[gerrit:1342118{{!}}Set up instrument for 5-arm test (T437076)]]
* 06:50 moritzm: installing sudo security updates
* 06:47 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply
* 06:37 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply
* 05:12 moritzm: pruned obsolete Bullseye image buildkitd from the docker registry [[phab:T416452|T416452]]
* 04:56 kart_: Updated Apertium to 2026-09-15-084320-production ([[phab:T437213|T437213]])
* 04:54 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/apertium: apply
* 04:54 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/apertium: apply
* 04:50 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/apertium: apply
* 04:49 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/apertium: apply
* 04:45 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/apertium: apply
* 04:45 kartik@deploy1003: helmfile [staging] START helmfile.d/services/apertium: apply
* 04:24 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 04:24 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 04:22 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 04:22 oblivian@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 36s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-15 ==
* 23:09 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply
* 23:08 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply
* 23:08 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply
* 23:08 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply
* 23:07 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 23:07 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 23:07 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 23:07 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 23:06 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 23:06 rzl@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 22:57 sukhe@puppetserver1001: conftool action : set/weight=1; selector: name=cp6001.drmrs.wmnet,service=cdn
* 22:50 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc-wf1002.eqiad.wmnet with OS trixie
* 22:46 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs ([[phab:T436363|T436363]])
* 22:46 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs6003.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:46 brett@cumin2003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs6003.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:46 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs6003.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:46 brett@cumin2003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs6003.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:45 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs6002.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:45 brett@cumin2003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs6002.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:45 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs6002.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:44 brett@cumin2003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs6002.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:44 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs6001.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:44 brett@cumin2003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs6001.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:43 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs6001.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:43 brett@cumin2003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs6001.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:43 brett@cumin2003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs ([[phab:T436363|T436363]])
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc-wf1002.eqiad.wmnet with reason: host reimage
* 22:33 brett@puppetserver1001: conftool action : set/weight=100; selector: name=cp6001.*
* 22:32 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp6001.*
* 22:26 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc-wf1002.eqiad.wmnet with reason: host reimage
* 22:07 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc-wf1002
* 22:07 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-wf1002
* 22:07 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host mc-wf1002
* 22:07 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc-wf1002.eqiad.wmnet 142.48.64.10.in-addr.arpa 2.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:07 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache mc-wf1002.eqiad.wmnet 142.48.64.10.in-addr.arpa 2.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:07 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:07 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc-wf1002 - rzl@cumin2003"
* 22:07 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc-wf1002 - rzl@cumin2003"
* 22:02 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:01 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host mc-wf1002
* 22:01 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host mc-wf1002.eqiad.wmnet with OS trixie
* 21:57 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6001.drmrs.wmnet with OS trixie
* 21:55 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply
* 21:55 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply
* 21:53 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply
* 21:53 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply
* 21:53 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:53 rzl@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:52 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:52 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:48 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:48 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:34 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6001.drmrs.wmnet with reason: host reimage
* 21:30 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp6001.drmrs.wmnet with reason: host reimage
* 21:20 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342049{{!}}MobileFrontend: Add app icons (T434258)]] (duration: 11m 47s)
* 21:15 jdlrobson@deploy1003: jdlrobson, cklimas: Continuing with deployment
* 21:13 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp6001.drmrs.wmnet with OS trixie
* 21:12 jdlrobson@deploy1003: jdlrobson, cklimas: Backport for [[gerrit:1342049{{!}}MobileFrontend: Add app icons (T434258)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:12 brett@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp6001.mgmt.drmrs.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:08 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1342049{{!}}MobileFrontend: Add app icons (T434258)]]
* 20:51 cdobbins@puppetserver1001: conftool action : set/weight=1; selector: name=cp2046.codfw.wmnet
* 20:51 cdobbins@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp2046.codfw.wmnet
* 20:48 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341271{{!}}Parsoid Read Views: Enable on all namespaces on wikitech (labswiki) (T437916)]] (duration: 09m 11s)
* 20:47 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp2046.codfw.wmnet with OS trixie
* 20:43 arlolra@deploy1003: ssastry, arlolra: Continuing with deployment
* 20:42 arlolra@deploy1003: ssastry, arlolra: Backport for [[gerrit:1341271{{!}}Parsoid Read Views: Enable on all namespaces on wikitech (labswiki) (T437916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:38 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1341271{{!}}Parsoid Read Views: Enable on all namespaces on wikitech (labswiki) (T437916)]]
* 20:34 brett@cumin2003: START - Cookbook sre.hosts.provision for host cp6001.mgmt.drmrs.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:30 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp6001.*
* 20:24 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp2046.codfw.wmnet with reason: host reimage
* 20:23 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339741{{!}}Enable ReaderExperiments in eswiki, jawiki, and ptwiki (T438009)]] (duration: 15m 58s)
* 20:20 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cp2046.codfw.wmnet with reason: host reimage
* 20:19 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T436363|T436363]])
* 20:19 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs4010.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:19 brett@cumin2003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs4010.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:19 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs4010.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:19 brett@cumin2003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs4010.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:18 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs4009.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:18 arlolra@deploy1003: lwatson, arlolra: Continuing with deployment
* 20:18 brett@cumin2003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs4009.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:17 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs4009.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:17 brett@cumin2003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs4009.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:17 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs4008.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:17 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp3075.esams.wmnet
* 20:17 sukhe@puppetserver1001: conftool action : set/weight=1; selector: name=cp3075.esams.wmnet
* 20:17 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp1103.eqiad.wmnet
* 20:17 brett@cumin2003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs4008.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:16 sukhe@puppetserver1001: conftool action : set/weight=1; selector: name=cp1103.eqiad.wmnet
* 20:16 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs4008.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:16 brett@cumin2003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs4008.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:16 brett@cumin2003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T436363|T436363]])
* 20:11 arlolra@deploy1003: lwatson, arlolra: Backport for [[gerrit:1339741{{!}}Enable ReaderExperiments in eswiki, jawiki, and ptwiki (T438009)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:10 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) config_reloading A:liberica-ulsfo ([[phab:T436363|T436363]])
* 20:08 brett@cumin2003: START - Cookbook sre.loadbalancer.admin config_reloading A:liberica-ulsfo ([[phab:T436363|T436363]])
* 20:08 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp1103.eqiad.wmnet with OS trixie
* 20:07 inflatador: bking@ganeti1046 sudo gnt-instance modify -B memory=4g,vcpus=4 dse-k8s-etcd100[1-3].eqiad.wmnet [[phab:T438084|T438084]]
* 20:07 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1339741{{!}}Enable ReaderExperiments in eswiki, jawiki, and ptwiki (T438009)]]
* 20:06 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp3075.esams.wmnet with OS trixie
* 20:04 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp7009.*
* 20:04 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host cp2046.codfw.wmnet with OS trixie
* 20:02 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp7009.*
* 20:02 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp7009.*
* 20:02 brett@puppetserver1001: conftool action : set/weight=1; selector: name=cp7009.*
* 20:01 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp7009.magru.wmnet with OS trixie
* 19:53 brett@puppetserver1001: conftool action : set/weight=1; selector: name=cp4045.*
* 19:53 brett@puppetserver1001: conftool action : set/weight=1; selector: name=cp4046.*
* 19:52 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4046.*
* 19:51 cdobbins@puppetserver1001: conftool action : set/pooled=no; selector: name=cp2046.codfw.wmnet
* 19:51 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp4046.ulsfo.wmnet with OS trixie
* 19:49 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4045.*
* 19:45 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp1103.eqiad.wmnet with reason: host reimage
* 19:43 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp4045.ulsfo.wmnet with OS trixie
* 19:41 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp3075.esams.wmnet with reason: host reimage
* 19:39 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cp1103.eqiad.wmnet with reason: host reimage
* 19:37 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp7009.magru.wmnet with reason: host reimage
* 19:33 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cp3075.esams.wmnet with reason: host reimage
* 19:32 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp7009.magru.wmnet with reason: host reimage
* 19:27 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp4046.ulsfo.wmnet with reason: host reimage
* 19:23 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp4046.ulsfo.wmnet with reason: host reimage
* 19:21 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp1103.eqiad.wmnet with OS trixie
* 19:19 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp1103.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:19 sukhe@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cp1103.eqiad.wmnet with reason: reimage
* 19:18 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp4045.ulsfo.wmnet with reason: host reimage
* 19:13 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 19:12 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 19:12 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 19:12 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp3075.esams.wmnet with OS trixie
* 19:12 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp4045.ulsfo.wmnet with reason: host reimage
* 19:12 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 19:11 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp3075.mgmt.esams.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:10 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp7009.magru.wmnet with OS trixie
* 19:09 brett@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp7009.mgmt.magru.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:08 sukhe@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cp3075.esams.wmnet with reason: reimaging
* 19:08 sukhe@cumin1004: START - Cookbook sre.hosts.provision for host cp1103.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:05 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp4046.ulsfo.wmnet with OS trixie
* 19:05 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore1006.eqiad.wmnet with OS bookworm
* 19:04 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3075.*
* 19:02 brett@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp4046.mgmt.ulsfo.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:01 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp1103.eqiad.wmnet
* 19:01 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp1103.eqiad.wmnet
* 19:00 sukhe@cumin1004: START - Cookbook sre.hosts.provision for host cp3075.mgmt.esams.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:58 brett@cumin2003: START - Cookbook sre.hosts.provision for host cp7009.mgmt.magru.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:57 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3075.esams.wmnet
* 18:55 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp1101.eqiad.wmnet
* 18:55 sukhe@puppetserver1001: conftool action : set/weight=1; selector: name=cp1101.eqiad.wmnet
* 18:55 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp4045.ulsfo.wmnet with OS trixie
* 18:52 brett@cumin2003: START - Cookbook sre.hosts.provision for host cp4046.mgmt.ulsfo.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:52 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp1101.eqiad.wmnet with OS trixie
* 18:45 cdobbins@puppetserver1001: conftool action : set/weight=1; selector: name=cp2044.codfw.wmnet
* 18:44 cdobbins@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp2044.codfw.wmnet
* 18:44 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore1006.eqiad.wmnet with reason: host reimage
* 18:40 brett@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp4045.mgmt.ulsfo.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:40 eevans@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore1006.eqiad.wmnet with reason: host reimage
* 18:39 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp3074.esams.wmnet
* 18:36 sukhe@cumin1004: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) config_reloading P<nowiki>{</nowiki>lvs3008.esams.wmnet<nowiki>}</nowiki> and A:liberica
* 18:36 sukhe@cumin1004: START - Cookbook sre.loadbalancer.admin config_reloading P<nowiki>{</nowiki>lvs3008.esams.wmnet<nowiki>}</nowiki> and A:liberica
* 18:33 sukhe@puppetserver1001: conftool action : set/weight=1; selector: name=cp3074.esams.wmnet
* 18:32 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp2044.codfw.wmnet with OS trixie
* 18:31 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp3074.esams.wmnet with OS trixie
* 18:30 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp1101.eqiad.wmnet with reason: host reimage
* 18:29 brett@cumin2003: START - Cookbook sre.hosts.provision for host cp4045.mgmt.ulsfo.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:26 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cp1101.eqiad.wmnet with reason: host reimage
* 18:22 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp7009.magru.wmnet with OS trixie
* 18:20 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1006.eqiad.wmnet with OS bookworm
* 18:20 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:19 eevans@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:19 eevans@cumin1004: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore1006.eqiad.wmnet
* 18:19 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore1006.eqiad.wmnet
* 18:10 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp1101.eqiad.wmnet with OS trixie
* 18:09 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp2044.codfw.wmnet with reason: host reimage
* 18:09 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp4045.ulsfo.wmnet with OS trixie
* 18:08 eevans@cumin1004: START - Cookbook sre.hosts.reboot-single for host sessionstore1006.eqiad.wmnet
* 18:07 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp3074.esams.wmnet with reason: host reimage
* 17:52 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp7009.magru.wmnet with reason: host reimage
* 17:48 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp2044.codfw.wmnet with OS trixie
* 17:43 cdobbins@puppetserver1001: conftool action : set/pooled=no; selector: name=cp2044.codfw.wmnet
* 17:36 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp3074.esams.wmnet with OS trixie
* 17:34 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp4046.ulsfo.wmnet with OS trixie
* 17:34 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp4045.ulsfo.wmnet with OS trixie
* 17:33 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3074.esams.wmnet
* 17:28 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp4046.*
* 17:28 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp4045.*
* 17:26 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp7009.magru.wmnet with OS trixie
* 17:25 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp1101.eqiad.wmnet with OS trixie
* 17:24 cdobbins@puppetserver1001: conftool action : set/pooled=no; selector: name=cp7009.magru.wmnet
* 17:23 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp1101.eqiad.wmnet
* 17:22 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:22 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org
* 17:02 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp7009.*
* 17:01 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:01 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:00 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:59 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:59 vriley@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host ms-be1099
* 16:59 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ms-be1099
* 16:59 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:56 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 16:55 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:55 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:55 vriley@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host ms-be1099
* 16:55 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ms-be1099
* 16:55 vriley@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host ms-be1099
* 16:54 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ms-be1099
* 16:54 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1098.eqiad.wmnet with OS bullseye
* 16:53 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:53 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [ms-be1099] - vriley@cumin1003"
* 16:53 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [ms-be1099] - vriley@cumin1003"
* 16:49 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 16:33 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS bullseye
* 16:21 mutante: temp disabling puppet on C:zookeeper (32 hosts) - safe deploy of https://gerrit.wikimedia.org/r/c/operations/puppet/+/1327569
* 16:04 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341890{{!}}Restore table borders for client-side MathJax (T435274)]], [[gerrit:1340558{{!}}lift IP cap for edit-a-thon /workshop (T437609 T437594 T437470)]] (duration: 24m 19s)
* 15:59 krinkle@deploy1003: anzx, krinkle: Continuing with deployment
* 15:44 krinkle@deploy1003: anzx, krinkle: Backport for [[gerrit:1341890{{!}}Restore table borders for client-side MathJax (T435274)]], [[gerrit:1340558{{!}}lift IP cap for edit-a-thon /workshop (T437609 T437594 T437470)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:40 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1341890{{!}}Restore table borders for client-side MathJax (T435274)]], [[gerrit:1340558{{!}}lift IP cap for edit-a-thon /workshop (T437609 T437594 T437470)]]
* 15:34 elukey@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 15:34 elukey@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 15:33 brennen@deploy1003: Finished deploy [phabricator/deployment@c386249]: deploy phab1005 for [[phab:T437930|T437930]] (duration: 00m 39s)
* 15:33 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS bullseye
* 15:33 brennen@deploy1003: Started deploy [phabricator/deployment@c386249]: deploy phab1005 for [[phab:T437930|T437930]]
* 15:32 elukey@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 15:32 elukey@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'sync'.
* 15:32 brennen@deploy1003: Finished deploy [phabricator/deployment@c386249]: deploy phab2003 for [[phab:T437930|T437930]] (duration: 00m 52s)
* 15:32 elukey@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 15:32 elukey@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'sync'.
* 15:31 brennen@deploy1003: Started deploy [phabricator/deployment@c386249]: deploy phab2003 for [[phab:T437930|T437930]]
* 15:31 elukey@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'.
* 15:31 elukey@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'.
* 15:26 jelto@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on phab2003.codfw.wmnet,phab[1005-1006].eqiad.wmnet with reason: Phabricator deploy
* 15:26 moritzm: pruned obsolete Bullseye image amd-gpu-tester from the docker registry [[phab:T416452|T416452]]
* 15:12 elukey@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 15:12 elukey@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 15:11 elukey@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 15:11 elukey@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 15:00 tgr@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341292{{!}}CommonSettings: Use a restrictive CSP for auth.wikimedia.org (T419684)]] (duration: 25m 11s)
* 14:55 tgr@deploy1003: tgr, arendpieter: Continuing with deployment
* 14:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:53 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:53 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:52 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:49 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 14:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 14:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 14:47 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 14:47 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 14:45 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 14:45 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 14:45 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 14:44 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 14:42 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 14:42 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 14:39 tgr@deploy1003: tgr, arendpieter: Backport for [[gerrit:1341292{{!}}CommonSettings: Use a restrictive CSP for auth.wikimedia.org (T419684)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:34 tgr@deploy1003: Started scap sync-world: Backport for [[gerrit:1341292{{!}}CommonSettings: Use a restrictive CSP for auth.wikimedia.org (T419684)]]
* 14:17 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341895{{!}}ReportIncidentController: Instance cache expensive methods (T437588)]] (duration: 11m 56s)
* 14:16 btullis@cumin1004: END (PASS) - Cookbook sre.ceph.rotate-osd-keys (exit_code=0) rolling rotate_keys on A:cephosd-codfw
* 14:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:12 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 14:09 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1341895{{!}}ReportIncidentController: Instance cache expensive methods (T437588)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:05 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1341895{{!}}ReportIncidentController: Instance cache expensive methods (T437588)]]
* 13:43 btullis@cumin1004: START - Cookbook sre.ceph.rotate-osd-keys rolling rotate_keys on A:cephosd-codfw
* 13:36 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341861{{!}}SuggestedInvestigations: Update "sockpuppet" queue view defaults (T438018)]] (duration: 10m 23s)
* 13:32 stran@deploy1003: stran: Continuing with deployment
* 13:30 stran@deploy1003: stran: Backport for [[gerrit:1341861{{!}}SuggestedInvestigations: Update "sockpuppet" queue view defaults (T438018)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:26 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1341861{{!}}SuggestedInvestigations: Update "sockpuppet" queue view defaults (T438018)]]
* 13:21 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:20 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:19 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334946{{!}}ArticleGuidance: Remove the experiment configuration keys (T434487)]] (duration: 09m 19s)
* 13:16 btullis@cumin1004: END (PASS) - Cookbook sre.ceph.rotate-osd-keys (exit_code=0) rolling rotate_keys on P<nowiki>{</nowiki>cephosd2001.codfw.wmnet<nowiki>}</nowiki> and (A:cephosd-codfw or A:cephosd-eqiad)
* 13:15 sbisson@deploy1003: sbisson: Continuing with deployment
* 13:14 sbisson@deploy1003: sbisson: Backport for [[gerrit:1334946{{!}}ArticleGuidance: Remove the experiment configuration keys (T434487)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) config_reloading P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:10 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1334946{{!}}ArticleGuidance: Remove the experiment configuration keys (T434487)]]
* 13:10 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin config_reloading P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:08 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:08 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:07 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:07 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:07 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:07 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:03 btullis@cumin1004: START - Cookbook sre.ceph.rotate-osd-keys rolling rotate_keys on P<nowiki>{</nowiki>cephosd2001.codfw.wmnet<nowiki>}</nowiki> and (A:cephosd-codfw or A:cephosd-eqiad)
* 13:00 btullis@cumin1004: END (PASS) - Cookbook sre.ceph.rotate-osd-keys (exit_code=0) rolling rotate_keys on P<nowiki>{</nowiki>cephosd2001.codfw.wmnet<nowiki>}</nowiki> and (A:cephosd-codfw or A:cephosd-eqiad)
* 12:59 btullis@cumin1004: START - Cookbook sre.ceph.rotate-osd-keys rolling rotate_keys on P<nowiki>{</nowiki>cephosd2001.codfw.wmnet<nowiki>}</nowiki> and (A:cephosd-codfw or A:cephosd-eqiad)
* 12:46 btullis@cumin1004: END (PASS) - Cookbook sre.ceph.rotate-osd-keys (exit_code=0) rolling rotate_keys on P<nowiki>{</nowiki>cephosd2001.codfw.wmnet<nowiki>}</nowiki> and (A:cephosd-codfw or A:cephosd-eqiad)
* 12:45 btullis@cumin1004: START - Cookbook sre.ceph.rotate-osd-keys rolling rotate_keys on P<nowiki>{</nowiki>cephosd2001.codfw.wmnet<nowiki>}</nowiki> and (A:cephosd-codfw or A:cephosd-eqiad)
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: network maintenance finished, [[phab:T437984|T437984]]]
* 12:40 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: network maintenance finished, [[phab:T437984|T437984]]]
* 12:29 XioNoX: asw1-b4-magru> request system reboot - [[phab:T437984|T437984]]
* 12:24 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 12:24 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 12:24 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 12:23 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 12:23 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 12:23 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 12:22 moritzm: installing shadow security updates
* 12:19 slyngshede@puppetserver1001: conftool action : set/weight=1; selector: name=cp7010.magru.wmnet
* 12:13 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 12:13 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 12 hosts with reason: Switch maintenance
* 12:12 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-b4-magru,asw1-b4-magru IPv6,asw1-b4-magru.mgmt with reason: Switch maintenance
* 12:11 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 12:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: switch reboot, [[phab:T437984|T437984]]]
* 12:11 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: switch reboot, [[phab:T437984|T437984]]]
* 12:09 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on install7002.wikimedia.org with reason: switch reboot
* 12:08 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 12:07 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 12:07 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 12:07 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 12:06 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 12:06 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 12:03 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 12:03 XioNoX: push pfw policies - [[phab:T437627|T437627]]
* 12:01 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 12:01 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) config_reloading P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 12:00 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin config_reloading P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 11:56 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 11:56 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 11:33 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2250.codfw.wmnet with reason: cloning db2201
* 11:18 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7004.magru.wmnet
* 11:17 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7004.magru.wmnet
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 11:05 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 11:05 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 11:04 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 10:51 slyngshede@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp7010.magru.wmnet
* 10:34 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 10:33 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 10:30 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1098.eqiad.wmnet with OS trixie
* 10:21 moritzm: failover Ganeti master in magru to ganeti7001
* 10:20 moritzm: increased DRBD replication speed in Ganeti/magru [[phab:T428878|T428878]]
* 10:10 hashar@deploy1003: Finished deploy [integration/docroot@5cf09c8]: build: Updating npm dependencies (duration: 00m 13s)
* 10:10 hashar@deploy1003: Started deploy [integration/docroot@5cf09c8]: build: Updating npm dependencies
* 10:09 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 10:08 moritzm: increased DRBD replication speed in Ganeti/esams [[phab:T428878|T428878]]
* 10:07 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 10:05 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 10:05 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:39 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: switches reboot, [[phab:T437984|T437984]]]
* 09:39 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool esams [reason: switches reboot, [[phab:T437984|T437984]]]
* 09:31 XioNoX: asw1-by27-esams> request system reboot - [[phab:T437984|T437984]]
* 09:30 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS trixie
* 09:28 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp7010.magru.wmnet with OS trixie
* 09:26 ayounsi@cumin1003: END (FAIL) - Cookbook sre.network.depool-rack (exit_code=99) with action 'depool' for esams rack BY27
* 09:24 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for esams rack BY27
* 09:24 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1098.eqiad.wmnet with OS trixie
* 09:23 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS trixie
* 09:22 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be1098.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 09:15 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.20 refs [[phab:T430839|T430839]]
* 09:10 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1098.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 09:06 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-by27-esams,asw1-by27-esams IPv6,asw1-by27-esams.mgmt with reason: Switch maintenance
* 09:05 ayounsi@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on asw1-by27-esams IPv6,asw1-by27-esams.mgmt,asw1-by-27-esams with reason: Switch maintenance
* 09:04 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 12 hosts with reason: Switch maintenance
* 09:04 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp7010.magru.wmnet with reason: host reimage
* 09:01 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: switches reboot, [[phab:T437984|T437984]]]
* 09:00 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: switches reboot, [[phab:T437984|T437984]]]
* 09:00 slyngshede@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp7010.magru.wmnet with reason: host reimage
* 08:59 jnuche@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341698{{!}}RestSandbox: Pass JsonLocalizer instead of ResponseFactory to ModuleManager (T437982)]] (duration: 12m 03s)
* 08:55 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2197.codfw.wmnet with reason: cloning db2201
* 08:55 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on install3004.wikimedia.org with reason: switch reboot
* 08:53 jnuche@deploy1003: jnuche: Continuing with deployment
* 08:52 jnuche@deploy1003: jnuche: Backport for [[gerrit:1341698{{!}}RestSandbox: Pass JsonLocalizer instead of ResponseFactory to ModuleManager (T437982)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:49 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: sync
* 08:49 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: sync
* 08:47 jnuche@deploy1003: Started scap sync-world: Backport for [[gerrit:1341698{{!}}RestSandbox: Pass JsonLocalizer instead of ResponseFactory to ModuleManager (T437982)]]
* 08:33 slyngshede@cumin1003: START - Cookbook sre.hosts.reimage for host cp7010.magru.wmnet with OS trixie
* 08:26 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be1098.eqiad.wmnet with OS trixie
* 08:26 slyngshede@puppetserver1001: conftool action : set/pooled=no; selector: name=cp7010.magru.wmnet
* 08:25 XioNoX: asw1-b3-magru - Disable logging and file logging for BRCM_PKT - [[phab:T437984|T437984]]
* 08:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS trixie
* 08:18 dpogorzelski@dns1004: END - running authdns-update
* 08:15 dpogorzelski@dns1004: START - running authdns-update
* 08:14 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3005.esams.wmnet
* 08:13 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3005.esams.wmnet
* 08:07 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341274{{!}}SI: Implement "queue view" functionality (T437183)]], [[gerrit:1341242{{!}}SuggestedInvestigations: Add and enable 'sockpuppets' queue view (T437183)]], [[gerrit:1341254{{!}}Add wmf-specific Special:SuggestedInvestigations messages (T437183)]] (duration: 55m 27s)
* 07:54 stran@deploy1003: stran: Continuing with deployment
* 07:31 stran@deploy1003: stran: Backport for [[gerrit:1341274{{!}}SI: Implement "queue view" functionality (T437183)]], [[gerrit:1341242{{!}}SuggestedInvestigations: Add and enable 'sockpuppets' queue view (T437183)]], [[gerrit:1341254{{!}}Add wmf-specific Special:SuggestedInvestigations messages (T437183)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:18 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply
* 07:18 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply
* 07:11 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1341274{{!}}SI: Implement "queue view" functionality (T437183)]], [[gerrit:1341242{{!}}SuggestedInvestigations: Add and enable 'sockpuppets' queue view (T437183)]], [[gerrit:1341254{{!}}Add wmf-specific Special:SuggestedInvestigations messages (T437183)]]
* 07:06 moritzm: pruned obsolete Bullseye image python3-devel from the docker registry [[phab:T416452|T416452]]
* 06:51 moritzm: pruned obsolete Bullseye image python3-build-bullseye from the docker registry [[phab:T416452|T416452]]
* 05:59 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply
* 05:49 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply
* 05:48 oblivian@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 05:47 oblivian@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 05:38 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply
* 05:28 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-video: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-video: apply
* 05:10 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-video: apply
* 05:10 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-video: apply
* 05:10 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-video: apply
* 05:10 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-video: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-syntaxhighlight: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-syntaxhighlight: apply
* 05:10 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-syntaxhighlight: apply
* 05:10 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-syntaxhighlight: apply
* 05:10 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-syntaxhighlight: apply
* 05:10 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-syntaxhighlight: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-media: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-media: apply
* 05:10 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-media: apply
* 05:10 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-media: apply
* 05:10 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-media: apply
* 05:10 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-media: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-constraints: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-constraints: apply
* 05:10 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-constraints: apply
* 05:09 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-constraints: apply
* 05:09 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-constraints: apply
* 05:09 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-constraints: apply
* 05:08 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply
* 05:08 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox: apply
* 05:07 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply
* 05:07 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox: apply
* 05:07 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox: apply
* 05:07 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox: apply
* 05:07 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply
* 05:07 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply
* 05:06 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-timeline: apply
* 05:06 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-timeline: apply
* 05:06 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply
* 05:06 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply
* 04:07 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.17 (duration: 07m 10s)
* 03:06 mwpresync@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.19,1.47.0-wmf.20,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.20 refs [[phab:T430839|T430839]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 22s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:43 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 00:23 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore1005.eqiad.wmnet with reason: host reimage
* 00:19 jclark@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore1005.eqiad.wmnet with reason: host reimage
* 00:17 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 00:07 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sessionstore1005.eqiad.wmnet with OS bookworm
== 2026-09-14 ==
* 23:41 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 23:26 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324966{{!}}Enable Produnto on pilot wikis (T421436)]] (duration: 12m 59s)
* 23:22 tstarling@deploy1003: tstarling: Continuing with deployment
* 23:17 tstarling@deploy1003: tstarling: Backport for [[gerrit:1324966{{!}}Enable Produnto on pilot wikis (T421436)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:13 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1324966{{!}}Enable Produnto on pilot wikis (T421436)]]
* 23:01 eevans@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 22:41 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341385{{!}}VisualEditor: don't register settings tool in wikitextCommandRegistry (T437810)]] (duration: 09m 22s)
* 22:36 kemayo@deploy1003: kemayo: Continuing with deployment
* 22:36 kemayo@deploy1003: kemayo: Backport for [[gerrit:1341385{{!}}VisualEditor: don't register settings tool in wikitextCommandRegistry (T437810)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:31 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1341385{{!}}VisualEditor: don't register settings tool in wikitextCommandRegistry (T437810)]]
* 22:22 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 22:22 eevans@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 21:43 sbassett: Deployed security fix for [[phab:T435623|T435623]]
* 21:29 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1098.eqiad.wmnet with OS bullseye
* 21:29 sbassett: Deployed security fix for [[phab:T434372|T434372]]
* 21:26 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 21:26 eevans@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 21:05 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338999{{!}}Enable ReadingLists for all logged-in users on English Wikipedia (T434923)]], [[gerrit:1340004{{!}}Enable Reading Recommendations experiment on test wiki (T437665)]] (duration: 11m 03s)
* 21:00 aude@deploy1003: aude, jdlrobson: Continuing with deployment
* 20:58 aude@deploy1003: aude, jdlrobson: Backport for [[gerrit:1338999{{!}}Enable ReadingLists for all logged-in users on English Wikipedia (T434923)]], [[gerrit:1340004{{!}}Enable Reading Recommendations experiment on test wiki (T437665)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:54 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1338999{{!}}Enable ReadingLists for all logged-in users on English Wikipedia (T434923)]], [[gerrit:1340004{{!}}Enable Reading Recommendations experiment on test wiki (T437665)]]
* 20:47 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341278{{!}}[A11y] Add list semantics to ReadingList page (T435864 T434923)]] (duration: 12m 49s)
* 20:43 aude@deploy1003: aude, jdlrobson: Continuing with deployment
* 20:39 aude@deploy1003: aude, jdlrobson: Backport for [[gerrit:1341278{{!}}[A11y] Add list semantics to ReadingList page (T435864 T434923)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:34 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1341278{{!}}[A11y] Add list semantics to ReadingList page (T435864 T434923)]]
* 20:32 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339811{{!}}wmf-config: Register content/v2-beta REST module as disabled (T432798)]], [[gerrit:1338274{{!}}wikifunctions: Move abstract fragments to mainstash (T432849)]] (duration: 25m 21s)
* 20:28 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:27 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:27 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:27 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:25 jforrester@deploy1003: jforrester, aghirelli: Continuing with deployment
* 20:24 jforrester@deploy1003: jforrester, aghirelli: Backport for [[gerrit:1339811{{!}}wmf-config: Register content/v2-beta REST module as disabled (T432798)]], [[gerrit:1338274{{!}}wikifunctions: Move abstract fragments to mainstash (T432849)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS bullseye
* 20:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be1098.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 20:06 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1339811{{!}}wmf-config: Register content/v2-beta REST module as disabled (T432798)]], [[gerrit:1338274{{!}}wikifunctions: Move abstract fragments to mainstash (T432849)]]
* 20:04 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1098.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 20:03 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be1098
* 20:02 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ms-be1098
* 20:02 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 20:02 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [ms-be1098] - vriley@cumin1003"
* 20:02 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [ms-be1098] - vriley@cumin1003"
* 19:59 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 19:58 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 19:57 dzahn@dns1005: END - running authdns-update
* 19:55 eevans@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 19:55 dzahn@dns1005: START - running authdns-update
* 19:54 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 19:38 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341273{{!}}[CodeMirror] enable for new users (enwiki), new VE integration (global) (T288161 T432558)]] (duration: 33m 51s)
* 19:26 musikanimal@deploy1003: musikanimal: Continuing with deployment
* 19:22 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1341273{{!}}[CodeMirror] enable for new users (enwiki), new VE integration (global) (T288161 T432558)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:04 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1341273{{!}}[CodeMirror] enable for new users (enwiki), new VE integration (global) (T288161 T432558)]]
* 18:53 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:50 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 18:50 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 18:49 eevans@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 18:26 brett@cumin2003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d6-eqiad
* 18:26 brett@cumin2003: START - Cookbook sre.network.tls for network device lsw1-d6-eqiad
* 18:26 brett@cumin2003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-eqiad
* 18:26 brett@cumin2003: START - Cookbook sre.network.tls for network device ssw1-d8-eqiad
* 18:25 brett@cumin2003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-eqiad
* 18:25 brett@cumin2003: START - Cookbook sre.network.tls for network device lsw1-d4-eqiad
* 18:25 root@cumin2003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-eqiad
* 18:25 root@cumin2003: START - Cookbook sre.network.tls for network device lsw1-d2-eqiad
* 18:17 jclark@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:14 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337611{{!}}extension-list: Add ModeratorToolkit (T431000)]] (duration: 09m 34s)
* 18:10 samtar@deploy1003: samtar: Continuing with deployment
* 18:09 samtar@deploy1003: samtar: Backport for [[gerrit:1337611{{!}}extension-list: Add ModeratorToolkit (T431000)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:06 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 18:05 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1337611{{!}}extension-list: Add ModeratorToolkit (T431000)]]
* 18:01 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 17:48 jclark@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 17:07 tgr@deploy1003: Finished scap sync-world: Backport for [[gerrit:1330446{{!}}CommonSettings: Use a restrictive, eval-free CSP for auth.wikimedia.org (T419684)]] (duration: 23m 19s)
* 17:00 tgr@deploy1003: arendpieter, tgr: Rolling back deployment
* 16:52 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 16:51 eevans@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 16:49 tgr@deploy1003: arendpieter, tgr: Backport for [[gerrit:1330446{{!}}CommonSettings: Use a restrictive, eval-free CSP for auth.wikimedia.org (T419684)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:44 tgr@deploy1003: Started scap sync-world: Backport for [[gerrit:1330446{{!}}CommonSettings: Use a restrictive, eval-free CSP for auth.wikimedia.org (T419684)]]
* 16:09 Amir1: drop links tables from db1252 ([[phab:T437278|T437278]])
* 16:07 Amir1: drop links tables from db2240 ([[phab:T437278|T437278]])
* 16:05 Amir1: drop non-links tables from db2247 ([[phab:T437278|T437278]])
* 15:53 Lucas_WMDE: UTC afternoon backport+config window belatedly done
* 15:50 lucaswerkmeister-wmde@deploy1003: mwscript-k8s job started: namespaceDupes abstractwiki --fix # [[phab:T437772|T437772]]
* 15:49 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335401{{!}}Adjust extendedconfirmed calculation to first edit on viwiki (T437006)]], [[gerrit:1340505{{!}}core-Namespaces: Add AW and AWT alias for its talk in abstractwiki (T437772)]] (duration: 10m 23s)
* 15:48 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 15:47 eevans@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 15:45 lucaswerkmeister-wmde@deploy1003: bunnypranav, lucaswerkmeister-wmde, tryvix1509: Continuing with deployment
* 15:43 lucaswerkmeister-wmde@deploy1003: bunnypranav, lucaswerkmeister-wmde, tryvix1509: Backport for [[gerrit:1335401{{!}}Adjust extendedconfirmed calculation to first edit on viwiki (T437006)]], [[gerrit:1340505{{!}}core-Namespaces: Add AW and AWT alias for its talk in abstractwiki (T437772)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:39 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1335401{{!}}Adjust extendedconfirmed calculation to first edit on viwiki (T437006)]], [[gerrit:1340505{{!}}core-Namespaces: Add AW and AWT alias for its talk in abstractwiki (T437772)]]
* 15:36 elukey@dns1004: END - running authdns-update
* 15:36 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 15:35 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 15:35 moritzm: installing shadow security updates
* 15:35 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 15:34 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:33 elukey@dns1004: START - running authdns-update
* 15:33 eevans@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:32 eevans@cumin1004: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore1005.eqiad.wmnet
* 15:29 lucaswerkmeister-wmde@deploy1003: mwscript-k8s job started: namespaceDupes afwiki --fix # [[phab:T437576|T437576]]
* 15:29 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338902{{!}}afwiki: Create Draft and Draft talk namespaces (T437576)]] (duration: 15m 44s)
* 15:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 15:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 15:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 15:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 15:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 15:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 15:21 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, tryvix1509: Continuing with deployment
* 15:21 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:17 eevans@cumin1004: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore1005.eqiad.wmnet
* 15:17 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, tryvix1509: Backport for [[gerrit:1338902{{!}}afwiki: Create Draft and Draft talk namespaces (T437576)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:17 eevans@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore1005.eqiad.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 15:13 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore1004.eqiad.wmnet
* 15:13 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1338902{{!}}afwiki: Create Draft and Draft talk namespaces (T437576)]]
* 15:06 eevans@cumin1004: START - Cookbook sre.hosts.reboot-single for host sessionstore1004.eqiad.wmnet
* 14:59 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore1004.eqiad.wmnet with OS bookworm
* 14:38 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore1004.eqiad.wmnet with reason: host reimage
* 14:33 marostegui@dns1004: END - running authdns-update
* 14:32 eevans@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore1004.eqiad.wmnet with reason: host reimage
* 14:30 marostegui@dns1004: START - running authdns-update
* 14:15 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1004.eqiad.wmnet with OS bookworm
* 14:14 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore1004.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:13 eevans@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1004.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:13 eevans@cumin1004: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore1004.eqiad.wmnet
* 14:13 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore1004.eqiad.wmnet
* 14:05 moritzm: kick off a rebuild of base images on build2004
* 14:05 moritzm: kick off a rebuild of base images on build2004
* 14:00 eevans@cumin1004: START - Cookbook sre.hosts.reboot-single for host sessionstore1004.eqiad.wmnet
* 14:00 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:00 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 13:56 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:54 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:52 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:44 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 13:43 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 13:42 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:40 eevans@cumin1004: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore1004.eqiad.wmnet
* 13:40 eevans@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore1004.eqiad.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 13:40 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:40 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:39 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:39 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:39 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 13:38 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 13:38 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:38 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:38 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 13:37 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 13:37 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:36 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:36 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 13:35 oblivian@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:35 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 13:35 oblivian@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:35 oblivian@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:34 oblivian@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'.
* 13:34 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 13:34 oblivian@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:34 oblivian@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:34 oblivian@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:34 oblivian@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 13:34 oblivian@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'.
* 13:33 oblivian@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'.
* 13:33 oblivian@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
* 13:33 oblivian@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
* 13:33 oblivian@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 13:33 oblivian@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 13:33 oblivian@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 13:32 oblivian@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 13:32 oblivian@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 13:32 oblivian@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 13:32 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 13:32 oblivian@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 13:32 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 13:32 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 13:29 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:23 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1003.eqiad.wmnet
* 13:21 sukhe: sudo cumin -b11 "A:cp-text" "run-puppet-agent --enable 'merging CR 1338134'" [[phab:T425441|T425441]]
* 13:20 sukhe: sudo cumin -b11 "A:cp-text" "run-puppet-agent --enable 'merging CR 1338134'"[[phab:T425441|T425441]]
* 13:19 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:17 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1003.eqiad.wmnet
* 13:14 moritzm: installing Bird security updates
* 13:09 sukhe: sudo cumin "A:cp-text" "disable-puppet 'merging CR 1338134'"
* 13:06 jmm@dns1004: END - running authdns-update
* 13:04 jmm@dns1004: START - running authdns-update
* 12:58 moritzm: update Trixie installer image to 13.7 [[phab:T437715|T437715]]
* 12:58 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 12:54 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 12:52 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 12:49 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 12:48 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 12:47 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 12:44 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 12:44 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 12:42 oblivian@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 12:40 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 12:40 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 12:40 oblivian@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 12:39 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 12:39 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 12:39 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 12:37 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 12:37 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 12:36 marostegui@cumin1004: conftool action : set/pooled=yes; selector: name=clouddb1025.eqiad.wmnet,service=x4
* 12:34 _joe_: adding gvisor labels to all wikikube clusters nodes
* 12:30 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 12:14 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 12:14 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 11:22 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1260: After cloning
* 10:48 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:45 ladsgroup@dns1004: END - running authdns-update
* 10:42 ladsgroup@dns1004: START - running authdns-update
* 10:37 marostegui@cumin1004: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=x4
* 10:37 marostegui@cumin1004: START - Cookbook sre.mysql.pool pool db1260: After cloning
* 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 10:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 10:04 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 10:04 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 09:53 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:53 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:41 Amir1: drop links tables from db2172 ([[phab:T437278|T437278]])
* 09:40 Amir1: drop links tables from db1228 ([[phab:T437278|T437278]])
* 09:08 marostegui: Stop mariadb on db1260 to clone dbstore1007, there will be lag on wikireplicas:x4 https://phabricator.wikimedia.org/T437839
* 09:07 taavi@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1025.eqiad.wmnet
* 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1260: Needs to clone another host from this one
* 09:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1260: Needs to clone another host from this one
* 09:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb[1024-1025].eqiad.wmnet,db[1155,1260].eqiad.wmnet,dbstore1007.eqiad.wmnet with reason: Adding x4
* 08:44 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:42 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:41 moritzm: pruned obsolete Bullseye images php8.3-icu72-cli / php8.3-icu72-fpm-multiversion-base / php8.3-icu72-fpm from the docker registry [[phab:T416452|T416452]]
* 08:37 taavi@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet
* 08:37 taavi@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1024.eqiad.wmnet
* 08:36 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on dbstore1007.eqiad.wmnet with reason: Adding x4
* 08:35 moritzm: pruned obsolete Bullseye images php8.1-cli/php8.1-fpm/ php8.1-fpm-multiversion-base from the docker registry [[phab:T416452|T416452]]
* 08:10 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: sync
* 08:08 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: sync
* 07:58 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 23s)
* 07:57 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 07:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 07:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 07:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 07:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 07:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 07:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 07:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 07:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 07:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 07:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 07:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 07:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 07:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 07:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 07:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 07:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 07:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 07:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 07:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 07:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 07:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 07:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 07:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 07:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 07:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 07:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 07:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 07:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 07:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 07:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 07:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 07:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 07:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 07:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 07:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 07:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 07:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 07:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 07:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 07:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 07:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 07:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 07:31 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1340801{{!}}Allow title-like strings with Package: prefix in require() (T430644)]], [[gerrit:1340802{{!}}Runtime: Add a facility for loading files by title (T430644)]] (duration: 34m 30s)
* 07:18 tstarling@deploy1003: tstarling: Continuing with deployment
* 07:17 tstarling@deploy1003: tstarling: Backport for [[gerrit:1340801{{!}}Allow title-like strings with Package: prefix in require() (T430644)]], [[gerrit:1340802{{!}}Runtime: Add a facility for loading files by title (T430644)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 06:56 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1340801{{!}}Allow title-like strings with Package: prefix in require() (T430644)]], [[gerrit:1340802{{!}}Runtime: Add a facility for loading files by title (T430644)]]
* 06:26 TimStarling: on deploy1003: docker image pull docker-registry.wikimedia.org/php8.3-fpm-multiversion-base
* 05:51 _joe_: pulled bookworm:latest from build2004 to build2001 [[phab:T437829|T437829]]
* 05:39 _joe_: force-running build-base-images on build2004 for [[phab:T437829|T437829]]
* 04:53 TimStarling: on build2001 rebuilding base images [[phab:T437829|T437829]]
* 03:00 tstarling@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.18,1.47.0-wmf.19,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted
* 02:59 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1340801{{!}}Allow title-like strings with Package: prefix in require() (T430644)]], [[gerrit:1340802{{!}}Runtime: Add a facility for loading files by title (T430644)]]
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-13 ==
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 29s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-12 ==
* 19:40 ladsgroup@cumin1003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-eqiad
* 19:32 ladsgroup@cumin1003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-eqiad
* 19:30 ladsgroup@cumin1003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw
* 19:21 ladsgroup@cumin1003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 35s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-11 ==
* 21:51 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 21:50 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 16:47 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339808{{!}}Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813{{!}}Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]] (duration: 07m 23s)
* 16:43 sbassett@deploy1003: sbassett: Continuing with deployment
* 16:42 sbassett@deploy1003: sbassett: Backport for [[gerrit:1339808{{!}}Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813{{!}}Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:40 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1339808{{!}}Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813{{!}}Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]]
* 16:08 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 14:39 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 14:10 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:10 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:10 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:09 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 13:40 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:28 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:11 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:10 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: db1199 repool
* 11:05 moritzm: installing Linux 6.1.187 on Bookworm hosts
* 11:05 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Jcrespo out of all services on: 2443 hosts
* 10:44 aokoth@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2020 in turn
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2019 in turn
* 10:41 Emperor: restart versitygw@objectstorage0[0-3].service on backup2018 in turn
* 10:40 Emperor: restart versitygw@objectstorage0[0-3].service on backup2017 in turn
* 10:39 Emperor: restart versitygw@objectstorage0[0-3].service on backup2016 in turn
* 10:37 Emperor: restart versitygw@objectstorage0[0-3].service on backup2015 in turn
* 10:36 Emperor: restart versitygw@objectstorage0[0-3].service on backup1020 in turn
* 10:35 Emperor: restart versitygw@objectstorage0[0-3].service on backup1019 in turn
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1199: db1199 repool
* 10:33 Emperor: restart versitygw@objectstorage0[0-3].service on backup1018 in turn
* 10:32 Emperor: restart versitygw@objectstorage0[0-3].service on backup1017 in turn
* 10:30 Emperor: restart versitygw@objectstorage0[0-3].service on backup1016 in turn
* 10:20 Emperor: restart versitygw@objectstorage0[1-3].service on backup1015 in turn
* 10:17 Emperor: restart versitygw@objectstorage00.service on backup1015
* 10:15 aokoth@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 08:46 slyngs: Update CAS/SSO to CAS 7.3.8.3
* 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:45 slyngshede@dns1004: END - running authdns-update
* 08:43 slyngshede@dns1004: START - running authdns-update
* 08:36 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:28 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 7 hosts with reason: Restarting s5
* 08:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db[1154,1269].eqiad.wmnet with reason: Restarting s5
* 08:20 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:20 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Repooling db1159
* 07:58 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:58 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:54 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:28 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:27 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1199: Needs to clone another host from this one
* 07:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1199: Needs to clone another host from this one
* 07:16 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Repooling db1159
* 07:15 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1199.eqiad.wmnet with reason: Cloning s4
* 07:10 TimStarling: killed jobs for [[phab:T437056|T437056]] since they weren't purging
* 06:38 TimStarling: also started refreshLinks for ptwiki and zhwiki, reparsing ~3000 pages altogether [[phab:T437056|T437056]]
* 06:27 TimStarling: for [[phab:T437056|T437056]]: mwscript-k8s refreshLinks.php --wiki=eswiki --tracking-category scribunto-common-error-category
* 05:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1159: Needs to clone another host from this one
* 05:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1159: Needs to clone another host from this one
* 05:30 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1159.eqiad.wmnet with reason: Cloning
* 05:29 marostegui: Start cloning db1245:s5 [[phab:T437563|T437563]]
* 05:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet,db1245.eqiad.wmnet with reason: Cloning
* 05:25 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] (duration: 09m 59s)
* 05:21 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:20 tstarling@deploy1003: tstarling: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:15 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 50s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-10 ==
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1349.eqiad.wmnet
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1349.eqiad.wmnet
* 23:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1349.eqiad.wmnet
* 23:09 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1349
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1349
* 22:32 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1349
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:28 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1349
* 22:27 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:27 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1349.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1348.eqiad.wmnet
* 22:23 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 22:12 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] (duration: 10m 59s)
* 22:06 musikanimal@deploy1003: kemayo, musikanimal: Rolling back deployment
* 22:05 musikanimal@deploy1003: kemayo, musikanimal: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:01 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]]
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:00 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:52 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:47 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:47 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] (duration: 13m 23s)
* 21:46 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:46 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2006.codfw.wmnet
* 21:42 derenrich@deploy1003: derenrich: Continuing with deployment
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1348
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1348
* 21:39 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1348
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:39 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:37 derenrich@deploy1003: derenrich: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:35 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:34 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1348
* 21:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 21:33 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1348.eqiad.wmnet
* 21:33 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]]
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1348.eqiad.wmnet
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1348.eqiad.wmnet
* 21:31 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2006.codfw.wmnet
* 21:31 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] (duration: 09m 45s)
* 21:27 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:26 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]]
* 21:19 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:19 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2006.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 21:17 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] (duration: 13m 54s)
* 21:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2005.codfw.wmnet with OS bookworm
* 21:12 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:07 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdeb
* 21:03 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]]
* 20:54 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:52 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] (duration: 23m 49s)
* 20:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:47 jdrewniak@deploy1003: jdrewniak, milazg: Continuing with deployment
* 20:34 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1346.eqiad.wmnet
* 20:32 jdrewniak@deploy1003: jdrewniak, milazg: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2005.codfw.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:28 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]]
* 20:27 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2005.codfw.wmnet
* 20:26 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:24 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:22 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] (duration: 11m 24s)
* 20:17 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 20:15 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2005.codfw.wmnet
* 20:14 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:12 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 20:11 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:07 jdrewniak@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.18,1.47.0-wmf.19,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted
* 20:05 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:02 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:01 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2005.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 19:53 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:50 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1346
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 19:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1346.eqiad.wmnet
* 19:21 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 thcipriani: Gerrit downtime incoming for upgrade
* 19:17 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hadoop.roll-restart-workers (exit_code=0) restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 dzahn@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:30:00 on gerrit.wikimedia.org with reason: maintenance upgrade
* 19:16 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on gerrit2003.wikimedia.org with reason: maintenance upgrade
* 19:04 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 18:21 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e] (duration: 00m 59s)
* 18:20 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e]
* 18:19 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e] (duration: 05m 13s)
* 18:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 18:14 otto@deploy1003: Started deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e]
* 18:13 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e] (duration: 00m 39s)
* 18:13 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e]
* 18:13 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 18:12 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 18:11 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 18:11 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit2002.wikimedia.org with reason: maintenance upgrade
* 18:11 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:11 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 18:10 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit1003.wikimedia.org with reason: maintenance upgrade
* 18:09 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 18:08 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 18:06 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 16:40 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 16:35 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 16:33 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]] synced to the te
* 16:28 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 15:33 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4
* 15:04 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] (duration: 08m 08s)
* 15:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host clouddumps1001.wikimedia.org with OS bookworm
* 14:59 samtar@deploy1003: samtar: Continuing with deployment
* 14:58 samtar@deploy1003: samtar: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:56 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]]
* 14:40 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:07 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 13:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host clouddumps1001.wikimedia.org with OS bookworm
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:55 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:51 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:42 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:41 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:48 klausman@dns1004: END - running authdns-update
* 12:46 klausman@dns1004: START - running authdns-update
* 12:35 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 12:35 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 12:05 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning x4
* 12:01 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2005.codfw.wmnet
* 11:55 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2005.codfw.wmnet
* 11:54 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1144.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:52 cgoubert@dns1004: END - running authdns-update
* 11:49 cgoubert@dns1004: START - running authdns-update
* 11:31 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2004.codfw.wmnet
* 11:25 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2004.codfw.wmnet
* 11:24 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:16 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1200.eqiad.wmnet
* 11:16 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1200.eqiad.wmnet
* 11:04 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1200.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:04 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1199.eqiad.wmnet
* 11:03 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1199.eqiad.wmnet
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:42 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1199.eqiad.wmnet with reason: Upgrading RAID firmware
* 10:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1024.eqiad.wmnet with reason: Cloning x4
* 10:00 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 09:56 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 09:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1024.eqiad.wmnet
* 09:45 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 09:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 09:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:04 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:51 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:46 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning x4
* 08:43 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:34 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] (duration: 09m 56s)
* 08:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:30 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:29 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:28 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:24 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]]
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 08:09 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 07:58 XioNoX: netflow1004:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:56 XioNoX: netflow2005:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 07:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 07:03 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:59 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:43 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:42 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:39 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 06:38 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 06:37 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 05:04 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 05:03 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tool-server' for release 'main' .
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 38s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1345.eqiad.wmnet
* 00:11 dzahn@dns1004: END - running authdns-update
* 00:08 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 00:08 dzahn@dns1004: START - running authdns-update
== 2026-09-09 ==
* 23:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:29 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1345
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 23:28 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1345.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1344.eqiad.wmnet
* 23:22 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] (duration: 11m 15s)
* 23:18 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 23:16 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 23:11 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]]
* 22:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:51 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2004.codfw.wmnet with OS bookworm
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1344
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1344
* 22:37 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1344
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1344
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 22:32 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] (duration: 10m 21s)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1344.eqiad.wmnet
* 22:27 derenrich@deploy1003: derenrich, egardner: Continuing with deployment
* 22:26 derenrich@deploy1003: derenrich, egardner: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:24 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:22 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]]
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:19 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] (duration: 13m 40s)
* 22:16 derenrich@deploy1003: derenrich: Rolling back deployment
* 22:10 derenrich@deploy1003: derenrich: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:05 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]]
* 22:03 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:44 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:40 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:36 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1343
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1343
* 21:27 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1343
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:27 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:23 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:22 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1343
* 21:22 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] (duration: 12m 29s)
* 21:21 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 21:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1343.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1342.eqiad.wmnet
* 21:17 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:14 jforrester@deploy1003: jforrester: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:09 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]]
* 21:08 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:07 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 21:07 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:06 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2004.codfw.wmnet
* 20:59 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:57 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:56 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2004.codfw.wmnet
* 20:53 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:53 bking@cumin2003: END (ERROR) - Cookbook sre.presto.roll-restart-workers (exit_code=97) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:53 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:50 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* {{safesubst:SAL entry|1=20:45 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2}}
* 20:42 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* 20:41 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 20:40 sbassett@deploy1003: aranyap, sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:39 sbassett@deploy1003: aranyap, sbassett: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "Filter}}
* 20:35 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* {{safesubst:SAL entry|1=20:34 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "}}
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:28 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:24 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 20:24 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1342
* 20:23 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 20:23 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1342.eqiad.wmnet
* 20:23 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1342.eqiad.wmnet
* 20:22 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1342.eqiad.wmnet
* 20:15 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1027.eqiad.wmnet with OS bookworm
* 19:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:41 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1027.eqiad.wmnet with OS bookworm
* 19:36 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:28 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:26 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:19 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 19:19 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:06 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 19:05 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 19:03 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 18:59 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 18:23 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:06 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1026.eqiad.wmnet with OS bookworm
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 18:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 17:54 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 17:49 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 17:45 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1026.eqiad.wmnet with OS bookworm
* 17:31 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:29 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 17:27 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 17:12 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 17:04 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 16:46 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] (duration: 09m 28s)
* 16:41 urbanecm@deploy1003: migr, urbanecm: Continuing with deployment
* 16:41 urbanecm@deploy1003: migr, urbanecm: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:36 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]]
* 16:28 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:25 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 15:55 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 15:54 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 15:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] (duration: 09m 43s)
* 15:41 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 15:40 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:36 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]]
* 15:17 urbanecm: Delete all running periodic jobs starting with `growthexperiments-refreshlinkrecommendations-*` (to pick up new configuration; [[phab:T392944|T392944]])
* 15:08 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:06 moritzm: installing grub2 bugfix updates from Bookworm point release
* 15:04 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp6008.drmrs.wmnet
* 15:01 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:42 hnowlan: half concurrency for parsoidCachePrewarm RecordLintJob and refreshLinks in jobqueue, eqiad & codfw
* 14:35 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 14:34 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:32 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 14:20 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:07 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2013.codfw.wmnet with OS trixie
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 14:06 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 13:55 btullis@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'.
* 13:53 btullis@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'.
* 13:43 btullis@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 13:42 btullis@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 13:29 moritzm: pruned obsolete Bullseye image dispatch from the docker registry [[phab:T416452|T416452]]
* 13:28 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:26 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b7-eqiad
* 13:25 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 13:24 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:22 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 13:22 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:17 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a4-eqiad
* 13:17 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] (duration: 10m 15s)
* 13:14 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 13:11 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 13:10 sbisson@deploy1003: sbisson: Continuing with deployment
* 13:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:04 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]]
* 13:02 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on ldap-rw[1001,2001].wikimedia.org with reason: work in progress
* 12:49 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 12:48 btullis@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
* 12:46 btullis@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
* 12:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] (duration: 14m 39s)
* 12:41 ladsgroup@deploy1003: tryvix1509, ladsgroup: Continuing with deployment
* 12:35 ladsgroup@deploy1003: tryvix1509, ladsgroup: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:31 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]]
* 12:16 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 12:16 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 11:53 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] (duration: 21m 58s)
* 11:48 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 11:35 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:31 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]]
* 10:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Repooling db2196
* 10:47 moritzm: pruned obsolete Bullseye images nodejs12-slim/nodejs12-devel/nodejs14-slim/nodejs16-slim from the docker registry [[phab:T416452|T416452]]
* 10:43 moritzm: installing Bird security updates
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1260: Repooling after cloning
* 10:09 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Repooling db2196
* 10:07 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 10:06 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 09:55 moritzm: pruned obsolete Bullseye images openjdk-8-jdk/openjdk-8-jre/openjdk-11-jre/openjdk-11-jdk from the docker registry [[phab:T416452|T416452]]
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1260: Repooling after cloning
* 09:52 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 09:52 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 09:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:28 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:27 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:03 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 09:02 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 09:01 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 16 hosts with reason: upgrade ssw1-a1-eqiad
* 08:58 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 22 hosts with reason: upgrade ssw1-a1-eqiad
* 08:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 08:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 08:49 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 08:49 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 08:48 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 08:48 cmooney@cumin1004: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 08:36 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning sanitarium
* 08:30 brouberol@dns1004: END - running authdns-update
* 08:28 moritzm: pruned obsolete Bullseye image golang1.15 from the docker registry [[phab:T416452|T416452]]
* 08:28 brouberol@dns1004: START - running authdns-update
* 08:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1260: Needs to clone another host from this one
* 08:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1260: Needs to clone another host from this one
* 08:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1260.eqiad.wmnet with reason: Cloning sanitarium
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:00 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 07:40 chlod: UTC morning backport window done
* 07:37 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] (duration: 21m 36s)
* 07:32 chlod@deploy1003: chlod, hamishz: Continuing with deployment
* 07:20 chlod@deploy1003: chlod, hamishz: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 45s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1025.eqiad.wmnet with OS bookworm
== 2026-09-08 ==
* 23:51 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:48 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:39 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1313.eqiad.wmnet
* 23:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:25 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 23:19 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:19 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:15 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:05 Amir1: dropped 57 tables on db1260 ([[phab:T437278|T437278]])
* 23:03 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 23:03 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:02 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 22:57 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1313
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1313
* 22:37 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1313
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:37 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:33 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 22:33 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1313
* 22:32 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 22:32 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1313.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1306.eqiad.wmnet
* 22:27 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6008.drmrs.wmnet with OS trixie
* 22:18 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 22:03 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 22:01 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] (duration: 09m 53s)
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 21:58 Amir1: drop links tables from db2210 ([[phab:T437278|T437278]])
* 21:57 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 21:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:56 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:52 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 21:51 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]]
* 21:51 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:48 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:45 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] (duration: 05m 27s)
* 21:43 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Continuing with deployment
* 21:40 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:39 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]]
* 21:35 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1024.eqiad.wmnet with OS bookworm
* 21:33 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp6008.drmrs.wmnet with OS trixie
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1306
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1306
* 21:30 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1306
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:24 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] (duration: 09m 12s)
* 21:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 21:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1306
* 21:19 reedy@deploy1003: reedy: Continuing with deployment
* 21:19 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:19 reedy@deploy1003: reedy: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1306.eqiad.wmnet
* 21:15 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:15 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]]
* {{safesubst:SAL entry|1=21:10 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out}}
* 21:05 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1024.eqiad.wmnet with OS bookworm
* 21:05 sbassett@deploy1003: sbassett: Continuing with deployment
* 21:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 21:04 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=21:03 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out non-http(s) lice}}
* 20:59 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=20:58 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out n}}
* 20:53 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 20:50 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1305.eqiad.wmnet
* 20:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 20:34 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 20:28 sbassett@deploy1003: sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:27 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-http(s) license}}
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1023.eqiad.wmnet with OS bookworm
* {{safesubst:SAL entry|1=20:23 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-}}
* 20:15 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] (duration: 10m 16s)
* 20:13 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:10 aaron@deploy1003: aaron: Continuing with deployment
* 20:09 aaron@deploy1003: aaron: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 20:05 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:05 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]]
* 20:01 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1023.eqiad.wmnet with OS bookworm
* 19:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1305
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1305
* 19:43 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1305
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:43 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:39 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 19:39 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1305
* 19:38 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 19:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1305.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1275.eqiad.wmnet
* 19:23 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 19:17 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 18:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:53 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1275
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1275
* 18:33 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1275
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1275
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 18:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1275.eqiad.wmnet
* 18:18 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 17:43 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy1003.eqiad.wmnet
* 17:36 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet
* 17:36 cdobbins@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-ntp (exit_code=0) rolling restart_daemons on A:dnsbox
* 17:30 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy1003.eqiad.wmnet
* 17:25 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet
* 17:15 swfrench@deploy1003: Finished scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup (duration: 04m 18s)
* 17:11 Amir1: dropping links tables from db1247 (s4 replica) - ([[phab:T437278|T437278]])
* 17:10 swfrench@deploy1003: Started scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup
* 16:51 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] (duration: 10m 19s)
* 16:46 zabe@deploy1003: zabe: Continuing with deployment
* 16:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]]
* 16:29 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 jhancock@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:12 jhancock@cumin2003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 16:08 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2003.codfw.wmnet
* 15:58 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:55 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2003.codfw.wmnet
* 15:54 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:44 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2002.codfw.wmnet
* 15:19 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2002.codfw.wmnet
* 14:55 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:46 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2001.codfw.wmnet
* 14:44 topranks: shutdown et-1/1/5 on cr1-codfw to shift traffic off ssw1-a1-codfw
* 14:43 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: upgrade ssw1-a1-eqiad
* 14:34 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2001.codfw.wmnet
* 14:33 btullis@cumin1003: END (ERROR) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=97) rolling reboot on A:cephosd-codfw
* 14:30 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet
* 14:28 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet
* 14:22 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet
* 14:13 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --olderThan {{Gerrit|1788220800}} --verbose # [[phab:T437158|T437158]]
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet
* 14:12 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet
* 14:03 topranks: drain traffic from ssw1-a1-codfw before JunOS upgrade [[phab:T426197|T426197]]
* 14:02 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet
* 13:58 cgoubert@deploy1003: helmfile [staging-codfw] DONE helmfile.d/services/mw-debug: apply
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet
* 13:57 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet
* 13:56 cgoubert@deploy1003: helmfile [staging-codfw] START helmfile.d/services/mw-debug: apply
* 13:55 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:54 stran@deploy1003: mwscript-k8s job started: foreachwikiindblist checkuser-suggested-investigations extensions/CheckUser/maintenance/populateSiCaseProperties.php # [[phab:T435066|T435066]]
* 13:52 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:51 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet
* 13:50 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:46 cdobbins@cumin1003: START - Cookbook sre.dns.roll-restart-ntp rolling restart_daemons on A:dnsbox
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet
* 13:46 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet
* 13:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2012.codfw.wmnet with OS bookworm
* 13:44 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] (duration: 34m 00s)
* 13:40 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet
* 13:37 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet
* 13:35 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet
* 13:32 stran@deploy1003: stran: Continuing with deployment
* 13:29 stran@deploy1003: stran: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet
* 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:21 moritzm: installing qemu security updates
* 13:18 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet
* 13:16 ayounsi@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'email' for AS: 139628
* 13:15 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 139628
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet
* 13:13 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet
* 13:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 2519
* 13:11 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 2519
* 13:10 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 14593
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]]
* 13:09 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:08 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 14593
* 13:06 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet
* 13:06 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2012.codfw.wmnet with OS bookworm
* 13:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2012.codfw.wmnet
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2012.codfw.wmnet
* 13:05 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:04 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet
* 13:01 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet
* 12:58 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 34655
* 12:58 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 34655
* 12:56 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2012.codfw.wmnet
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'clear' for AS: 35320
* 12:55 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'clear' for AS: 35320
* 12:55 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-codfw
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-codfw
* 12:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2011.codfw.wmnet with OS bookworm
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-by27-esams
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-by27-esams
* 12:54 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-esams
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-esams
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-esams
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-esams
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-eqsin
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-eqsin
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-ulsfo
* 12:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f3-eqiad
* 12:50 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f3-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-c8-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-c8-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:44 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:43 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet
* 12:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-eqiad
* 12:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f1-eqiad
* 12:41 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f1-eqiad
* 12:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-d5-eqiad
* 12:40 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-d5-eqiad
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet
* 12:38 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet
* 12:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:32 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet
* 12:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad
* 12:11 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2012.codfw.wmnet
* 12:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2011.codfw.wmnet with OS bookworm
* 12:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2011.codfw.wmnet
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2011.codfw.wmnet
* 12:09 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:06 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] (duration: 09m 54s)
* 12:01 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2011.codfw.wmnet
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Continuing with deployment
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2010.codfw.wmnet with OS bookworm
* 11:56 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]]
* 11:46 marostegui@dns1004: END - running authdns-update
* 11:44 marostegui@dns1004: START - running authdns-update
* 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:40 Amir1: dropping unneeded tables from x4 - db1260 ([[phab:T437278|T437278]])
* 11:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:23 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2011.codfw.wmnet
* 11:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2010.codfw.wmnet with OS bookworm
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2010.codfw.wmnet
* 11:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2010.codfw.wmnet
* 11:20 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2010.codfw.wmnet
* 11:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2009.codfw.wmnet with OS bookworm
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:43 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] (duration: 10m 57s)
* 10:39 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:38 samtar@deploy1003: samtar: Continuing with deployment
* 10:37 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:37 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:36 samtar@deploy1003: samtar: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:34 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:32 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]]
* 10:30 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2009.codfw.wmnet with OS bookworm
* 10:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2009.codfw.wmnet
* 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:18 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2009.codfw.wmnet
* 10:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2008.codfw.wmnet with OS bookworm
* 10:07 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:05 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:01 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:45 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] (duration: 13m 15s)
* 09:45 ayounsi@dns1004: END - running authdns-update
* 09:44 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:43 ayounsi@dns1004: START - running authdns-update
* 09:39 zabe@deploy1003: zabe: Continuing with deployment
* 09:37 zabe@deploy1003: zabe: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2008.codfw.wmnet with OS bookworm
* 09:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:32 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]]
* 09:32 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2008.codfw.wmnet
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2008.codfw.wmnet
* 09:31 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:29 XioNoX: remove GRE tunnels eqiad-drmrs eqdfw-ulsfo
* 09:23 moritzm: installing rsync security updates
* 09:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2007.codfw.wmnet with OS bookworm
* 09:22 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2008.codfw.wmnet
* 09:11 marostegui@cumin1003: dbctl commit (dc=all): 'Make x4 and s4 RW again [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96393 and previous config saved to /var/cache/conftool/dbconfig/20260908-091121-marostegui.json
* 09:07 marostegui@cumin1003: dbctl commit (dc=all): 'Remove old s4 masters from x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96392 and previous config saved to /var/cache/conftool/dbconfig/20260908-090749-marostegui.json
* 09:05 marostegui@cumin1003: dbctl commit (dc=all): 'Set x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96391 and previous config saved to /var/cache/conftool/dbconfig/20260908-090517-marostegui.json
* 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Set s4 commons to read-only for maintenance [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96389 and previous config saved to /var/cache/conftool/dbconfig/20260908-090228-marostegui.json
* 09:02 marostegui: Starting x4 split from s4, RO time on commons needed [[phab:T404715|T404715]]
* 09:00 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 08:58 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2008.codfw.wmnet
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:43 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 32 hosts with reason: x4 split
* 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2007.codfw.wmnet with OS bookworm
* 08:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:37 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:37 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2007.codfw.wmnet
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2007.codfw.wmnet
* 08:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:34 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:29 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:26 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2007.codfw.wmnet
* 08:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2006.codfw.wmnet with OS bookworm
* 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:50 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2007.codfw.wmnet
* 07:43 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2006.codfw.wmnet with OS bookworm
* 07:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2006.codfw.wmnet
* 07:37 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:30 denisse: Add grafana-plugins 0.15 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 07:29 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:27 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 07:27 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:27 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:22 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 07:22 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 07:18 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:14 jmm@dns1004: END - running authdns-update
* 07:13 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96388 and previous config saved to /var/cache/conftool/dbconfig/20260908-071308-marostegui.json
* 07:12 jmm@dns1004: START - running authdns-update
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96387 and previous config saved to /var/cache/conftool/dbconfig/20260908-071216-marostegui.json
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96386 and previous config saved to /var/cache/conftool/dbconfig/20260908-071159-marostegui.json
* 05:07 denisse: Add grafana-plugins 0.10 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.16 (duration: 02m 27s)
* 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]] (duration: 36m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 41s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-07 ==
* 21:52 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] (duration: 11m 00s)
* 21:47 zabe@deploy1003: zabe: Continuing with deployment
* 21:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]]
* 21:37 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] (duration: 09m 34s)
* 21:33 zabe@deploy1003: zabe: Continuing with deployment
* 21:32 zabe@deploy1003: zabe: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:28 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]]
* 21:03 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] (duration: 10m 27s)
* 20:58 zabe@deploy1003: zabe: Continuing with deployment
* 20:57 zabe@deploy1003: zabe: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:52 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]]
* 20:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set weight of db1261 to zero in s4 ([[phab:T437108|T437108]])', diff saved to https://phabricator.wikimedia.org/P96385 and previous config saved to /var/cache/conftool/dbconfig/20260907-203804-ladsgroup.json
* 20:23 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] (duration: 09m 28s)
* 20:19 zabe@deploy1003: zabe: Continuing with deployment
* 20:18 zabe@deploy1003: zabe: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:14 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]]
* 20:12 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] (duration: 10m 06s)
* 20:07 zabe@deploy1003: zabe, daimona: Continuing with deployment
* 20:06 zabe@deploy1003: zabe, daimona: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:02 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]]
* 19:59 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] (duration: 11m 27s)
* 19:55 zabe@deploy1003: zabe: Continuing with deployment
* 19:52 zabe@deploy1003: zabe: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:48 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]]
* 19:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] (duration: 11m 40s)
* 19:26 zabe@deploy1003: zabe: Continuing with deployment
* 19:23 zabe@deploy1003: zabe: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:19 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]]
* 19:07 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] (duration: 14m 17s)
* 19:00 zabe@deploy1003: zabe: Continuing with deployment
* 18:57 zabe@deploy1003: zabe: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:53 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]]
* 18:33 zabe@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 18:32 zabe@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 18:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] (duration: 09m 12s)
* 18:27 zabe@deploy1003: zabe: Continuing with deployment
* 18:26 zabe@deploy1003: zabe: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:22 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]]
* 16:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2005.codfw.wmnet with OS bookworm
* 16:01 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] (duration: 10m 22s)
* 15:59 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 15:57 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 15:56 zabe@deploy1003: zabe: Continuing with deployment
* 15:55 zabe@deploy1003: zabe: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:51 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]]
* 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:47 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 15:46 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 15:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:44 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 15:44 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 15:27 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2005.codfw.wmnet with OS bookworm
* 15:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2005.codfw.wmnet
* 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:11 moritzm: installing rsync security updates
* 15:04 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1341.eqiad.wmnet
* 15:03 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2004.codfw.wmnet with OS bookworm
* 15:00 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:58 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 14:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:40 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 14:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 14:35 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:32 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:29 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:25 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 14:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1228: Repooling db1228 into s4
* 14:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2004.codfw.wmnet with OS bookworm
* 14:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Repooling after cloning
* 14:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:19 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2004.codfw.wmnet
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2004.codfw.wmnet
* 14:18 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:14 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1074.eqiad.wmnet
* 14:14 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:13 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:13 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* {{safesubst:SAL entry|1=14:11 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mul}}
* 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2004.codfw.wmnet
* 14:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1074.eqiad.wmnet
* 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1073.eqiad.wmnet
* 14:07 krinkle@deploy1003: krinkle: Continuing with deployment
* {{safesubst:SAL entry|1=14:04 krinkle@deploy1003: krinkle: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with multiple properties}}
* 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1073.eqiad.wmnet
* 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1072.eqiad.wmnet
* 14:01 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1341
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1341
* 14:01 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 14:00 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1341
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* 13:59 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* {{safesubst:SAL entry|1=13:59 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mult}}
* 13:59 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2004.codfw.wmnet
* 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2003.codfw.wmnet with OS bookworm
* 13:55 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 13:55 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1072.eqiad.wmnet
* 13:55 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudvirt1067.eqiad.wmnet
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1341
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 13:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1341.eqiad.wmnet
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2008.wikimedia.org
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2008.wikimedia.org with OS trixie
* 13:39 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1067.eqiad.wmnet
* 13:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1066.eqiad.wmnet
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:37 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1228: Repooling db1228 into s4
* 13:36 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Repooling after cloning
* 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1066.eqiad.wmnet
* 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1065.eqiad.wmnet
* 13:28 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:27 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1065.eqiad.wmnet
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:25 moritzm: installing openssh security updates
* 13:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:24 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1340.eqiad.wmnet
* 13:23 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] (duration: 10m 06s)
* 13:17 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2003.codfw.wmnet with OS bookworm
* 13:17 stran@deploy1003: stran: Continuing with deployment
* 13:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:16 stran@deploy1003: stran: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:16 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:15 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 13:12 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]]
* 13:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2003.codfw.wmnet
* 13:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2003.codfw.wmnet
* 13:03 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2008.wikimedia.org with OS trixie
* 13:03 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:02 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:02 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:01 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:00 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2003.codfw.wmnet
* 12:58 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:54 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 12:54 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2008.wikimedia.org
* 12:47 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2003.codfw.wmnet
* 12:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2002.codfw.wmnet with OS bookworm
* 12:29 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 12:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2007.wikimedia.org
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2007.wikimedia.org with OS trixie
* 12:14 elukey: moved most of the Docker Registry's prefixes to a new internal S3 backend. For any docker pull failure that worked in the past, please ping me or drop a note in [[phab:T435499|T435499]] or contact the oncall SREs
* 12:07 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2002.codfw.wmnet with OS bookworm
* 12:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:02 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2002.codfw.wmnet
* 11:54 jmm@dns1004: END - running authdns-update
* 11:52 jmm@dns1004: START - running authdns-update
* 11:46 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2002.codfw.wmnet
* 11:46 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2007.wikimedia.org with OS trixie
* 11:46 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:46 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:41 moritzm: installing bash updates from bookworm point release
* 11:39 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:39 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2007.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts ldap-replica1006.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:35 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:32 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2001.codfw.wmnet with OS bookworm
* 11:28 jmm@cumin1004: START - Cookbook sre.hosts.decommission for hosts ldap-replica1006.wikimedia.org
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 11:18 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] (duration: 14m 08s)
* 11:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 11:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet
* 11:11 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:11 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:10 zabe@deploy1003: zabe: Continuing with deployment
* 11:10 zabe@deploy1003: zabe: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 11:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet
* 11:07 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet
* 11:07 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet
* 11:06 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:06 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:04 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]]
* 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 10:58 btullis@deploy1003: Finished scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]] (duration: 35m 20s)
* 10:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:53 jmm@dns1004: END - running authdns-update
* 10:51 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:50 jmm@dns1004: START - running authdns-update
* 10:47 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=frwiki # [[phab:T436659|T436659]]
* 10:40 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=hrwiki # [[phab:T436659|T436659]]
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:39 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 10:37 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 10:33 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 10:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2001.codfw.wmnet with OS bookworm
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1340
* 10:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1340.eqiad.wmnet
* 10:26 btullis@deploy1003: Started scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]]
* 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2001.codfw.wmnet
* 10:18 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2001.codfw.wmnet
* 10:13 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw
* 10:12 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:10 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:09 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:08 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:07 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:07 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:02 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:00 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:59 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:57 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/services/miscweb: apply
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/services/miscweb: apply
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:52 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:52 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] (duration: 10m 11s)
* 09:51 aokoth@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 09:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 09:49 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 09:49 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 09:48 aokoth@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 09:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:45 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:44 aokoth@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 09:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]]
* 09:41 aokoth@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 09:38 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Repooling after switchover
* 09:23 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw
* 09:15 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] (duration: 20m 12s)
* 09:12 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 139009
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1005.wikimedia.org
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1005.wikimedia.org with OS trixie
* 09:10 moritzm: rebuild software RAID following disk replacement [[phab:T437036|T437036]]
* 09:10 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 139009
* 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1022.eqiad.wmnet with OS bookworm
* 09:08 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 09:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:03 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet
* 09:02 moritzm: installing giflib security updates
* 09:01 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:59 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 08:57 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 08:56 jmm@cumin1004: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet
* 08:56 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:54 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]]
* 08:52 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:52 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:49 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 08:48 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:42 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:41 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:40 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:40 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96355 and previous config saved to /var/cache/conftool/dbconfig/20260907-083904-marostegui.json
* 08:38 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1005.wikimedia.org with OS trixie
* 08:38 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2157 to s5 primary [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96354 and previous config saved to /var/cache/conftool/dbconfig/20260907-083825-marostegui.json
* 08:38 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 marostegui: Starting s5 codfw failover from db2213 to db2157 - [[phab:T437188|T437188]]
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2157 with weight 0 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96353 and previous config saved to /var/cache/conftool/dbconfig/20260907-083448-marostegui.json
* 08:34 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s5 [[phab:T437188|T437188]]
* 08:28 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:28 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1005.wikimedia.org
* 08:22 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:20 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1022.eqiad.wmnet with OS bookworm
* 08:03 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:02 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 08:02 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:00 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 07:57 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 07:56 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1241.eqiad.wmnet with reason: Cloning
* 07:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1241: Cloning
* 07:55 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1241: Cloning
* 07:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Cloning
* 07:54 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Cloning
* 07:51 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 07:50 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 07:47 kartik@deploy1003: Finished scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] (duration: 41m 51s)
* 07:46 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 07:45 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 07:34 kartik@deploy1003: abi, kartik: Continuing with deployment
* 07:23 kartik@deploy1003: abi, kartik: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:05 kartik@deploy1003: Started scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]]
* 06:14 moritzm: installing Chromium security updates
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 08m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-06 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 25s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-05 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 26s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-04 ==
* 22:07 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 21:48 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:42 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:30 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:42 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:40 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 19:27 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:19 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:13 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:12 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host sretest2013
* 19:11 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host sretest2013
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:11 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:07 jhancock@cumin1003: START - Cookbook sre.dns.netbox
* 18:18 inflatador: bking@clouddumps100[12] `systemctl reset-failed` to quash alerts until https://w.wiki/UBje . The systemd timer should try again tomorrow
* 17:27 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b8-eqiad
* 17:27 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b8-eqiad
* 16:37 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:33 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-b7-eqiad
* 16:33 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 16:05 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b6-eqiad
* 16:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b6-eqiad
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1339.eqiad.wmnet
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 15:50 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:49 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:47 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:43 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 15:10 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b5-eqiad
* 15:09 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b5-eqiad
* 14:46 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1045.eqiad.wmnet
* 14:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1003.eqiad.wmnet with OS trixie
* 14:38 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 14:37 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b4-eqiad
* 14:37 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b4-eqiad
* 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339
* 14:32 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339
* 14:31 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 14:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet
* 14:26 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:24 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:21 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b3-eqiad
* 14:21 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b3-eqiad
* 14:17 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:17 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:04 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1003.eqiad.wmnet with OS trixie
* 13:54 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b2-eqiad
* 13:53 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b2-eqiad
* 13:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1002.eqiad.wmnet with OS trixie
* 13:18 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a4-eqiad
* 13:12 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:12 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b1-eqiad
* 13:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b1-eqiad
* 12:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1002.eqiad.wmnet with OS trixie
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow3004.esams.wmnet with OS trixie
* 12:33 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:28 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 12:15 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 12:11 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a7-eqiad
* 12:11 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a7-eqiad
* 12:01 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow3004.esams.wmnet with OS trixie
* 11:47 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a6-eqiad
* 11:47 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a6-eqiad
* 11:36 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki2003.codfw.wmnet
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki2003.codfw.wmnet with OS trixie
* 11:19 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:13 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a5-eqiad
* 11:06 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a5-eqiad
* 10:52 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki2003.codfw.wmnet with OS trixie
* 10:50 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:50 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:44 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 10:44 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki2003.codfw.wmnet
* 10:29 btullis@deploy1003: Finished scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli (duration: 41m 14s)
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 10:00 marostegui@cumin1003: Removing db1182 from zarcillo [[phab:T434869|T434869]]
* 10:00 marostegui@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db1182.eqiad.wmnet
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:57 btullis@deploy1003: Started scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli
* 09:53 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:53 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.decommission (exit_code=1)
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1182.eqiad.wmnet
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:50 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:46 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:45 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1182 from dbctl [[phab:T434869|T434869]]', diff saved to https://phabricator.wikimedia.org/P96346 and previous config saved to /var/cache/conftool/dbconfig/20260904-094527-marostegui.json
* 09:41 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki1003.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki1003.eqiad.wmnet with OS trixie
* 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1182: Decommissioning
* 09:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1182: Decommissioning
* 09:23 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 09:17 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow5003.eqsin.wmnet with OS trixie
* 09:02 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki1003.eqiad.wmnet with OS trixie
* 09:00 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 09:00 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:55 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:55 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki1003.eqiad.wmnet
* 08:54 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:48 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:45 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:40 btullis@deploy1003: Finished scap sync-world: Trying again for [[phab:T436913|T436913]] (duration: 34m 26s)
* 08:35 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2003.codfw.wmnet
* 08:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw2001.wikimedia.org with OS trixie
* 08:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4 days, 0:00:00 on db2196.codfw.wmnet with reason: Host crashed
* 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2003.codfw.wmnet
* 08:24 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:20 elukey@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:12 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:08 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:07 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 08:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:04 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 07:57 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow5003.eqsin.wmnet with OS trixie
* 07:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:55 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw2001.wikimedia.org with OS trixie
* 07:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:51 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:50 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie
* 07:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2196: down
* 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2196: down
* 06:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:52 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:40 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:28 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:21 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 06:20 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 39s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 01:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:02 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:59 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:46 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 00:44 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:39 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:23 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
== 2026-09-03 ==
* 21:46 tsev@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T435363|T435363]]
* 21:03 eevans@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:55 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] (duration: 12m 24s)
* 20:52 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:52 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1021.eqiad.wmnet with OS bookworm
* 20:50 arlolra@deploy1003: arlolra, tsev: Continuing with deployment
* 20:46 arlolra@deploy1003: arlolra, tsev: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:44 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1047.eqiad.wmnet
* 20:42 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]]
* 20:41 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a3-eqiad
* 20:40 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a3-eqiad
* 20:39 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] (duration: 10m 23s)
* 20:34 arlolra@deploy1003: arlolra, jgiannelos: Continuing with deployment
* 20:33 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1047.eqiad.wmnet
* 20:33 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:32 arlolra@deploy1003: arlolra, jgiannelos: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:31 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1046.eqiad.wmnet
* 20:28 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]]
* 20:23 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] (duration: 13m 41s)
* 20:20 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1046.eqiad.wmnet
* 20:16 catrope@deploy1003: catrope: Continuing with deployment
* 20:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1021.eqiad.wmnet with OS bookworm
* 20:13 catrope@deploy1003: catrope: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1021.eqiad.wmnet
* 20:09 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]]
* 20:00 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1021.eqiad.wmnet
* 19:49 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 19:19 swfrench@deploy1003: Finished scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]] (duration: 02m 59s)
* 19:16 swfrench@deploy1003: Started scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]]
* 19:01 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:01 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 18:57 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:57 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:34 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:20 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 17:55 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:55 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:53 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:52 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:46 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 17:46 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 17:45 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 17:44 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 17:44 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:40 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:39 ryankemper: [WDQS] Service looks healthy again, CPU load and thread count have dropped considerably over the last hour
* 17:39 ryankemper: [[phab:T421642|T421642]] [WDQS] requestctl changes: `2026-09-03 16:23-17:33` UTC: added hard-deny pair `cache-text/wdqs_futile_sparql_sep_2026_deny(+_bots)`; extended pattern `ua/wdqs_heavy_sparql_bots_2026` and added default-scope twin `wdqs_heavy_sparql_bots_jul_2026_ratelimit_default`; added ipblock `abuse/wdqs_sparql_scanners_sep_2026` + throttle `wdqs_sparql_scanners_sep_2026_ratelimit` (needed manual `requestctl update-provenance-map`)
* 17:37 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:37 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:31 andrew@cumin2003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudcephosd1045.eqiad.wmnet
* 17:30 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:30 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:28 dancy@deploy1003: Installation of scap version "4.289.0" completed for 3 hosts
* 17:26 dancy@deploy1003: Installing scap version "4.289.0" for 3 host(s)
* 17:24 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a2-eqiad
* 17:24 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:22 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:22 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a2-eqiad
* 17:14 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 17:10 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:10 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:09 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 17:01 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:59 andrew@cumin2003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 16:58 dancy: Running scap clean-images on deploy1003
* 16:52 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:50 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:39 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:14 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs2021.codfw.wmnet,service=wdqs-main
* 16:14 ryankemper: [[phab:T430880|T430880]] Stumbled across `wdqs2021` listed as inactive, looks like it was never fully re-pooled after a data xfer. Pooled.
* 16:12 btullis@deploy1003: Finished deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2] (duration: 00m 38s)
* 16:12 btullis@deploy1003: Started deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2]
* 16:12 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:07 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs101[1-4].eqiad.wmnet
* 16:03 btullis@deploy1003: Started scap sync-world: Rebuilding to pick up new version of dump scripts in mediawiki-cli for [[phab:T436913|T436913]]
* 16:01 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]] (duration: 09m 29s)
* 15:51 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=urldownloader[12]00[56].wikimedia.org [reason: depooling urldownloader trixie nodes]
* 15:51 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]]
* 15:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:29 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:24 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:24 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:18 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:15 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:09 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:05 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 15:04 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:02 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 15:00 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader
* 14:58 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw
* 14:53 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-ulsfo
* 14:49 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 14:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2002.codfw.wmnet
* 14:45 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:39 sukhe: sudo cumin "A:cp-text" "run-puppet-agent --enable 'merging CR 1334855'": [[phab:T425441|T425441]]
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2002.codfw.wmnet
* 14:32 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:30 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:29 sukhe: sudo cumin "A:cp-text" "disable-puppet 'merging CR 1334855'" [[phab:T425441|T425441]]
* 14:27 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-ulsfo
* 14:22 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:21 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:19 arnaudb@dns1006: END - running authdns-update
* 14:18 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1074
* 14:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1074
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-codfw
* 14:17 arnaudb@dns1006: START - running authdns-update
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:16 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:16 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1073
* 14:15 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1073
* 14:12 ayounsi@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host netflow2004.codfw.wmnet with OS trixie
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:10 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:08 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:08 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:07 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] (duration: 09m 36s)
* 14:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:05 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:03 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 14:03 samtar@deploy1003: btullis, samtar: Continuing with deployment
* 14:02 samtar@deploy1003: btullis, samtar: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:01 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:58 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]]
* 13:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:56 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:55 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:52 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqiad and A:durum
* 13:52 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-codfw
* 13:51 moritzm: installing sqlite3 security updates
* 13:51 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:51 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqiad and A:durum
* 13:49 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-codfw and A:durum
* 13:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:47 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-codfw and A:durum
* 13:47 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-esams
* 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:44 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:43 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:43 ayounsi@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:42 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] (duration: 13m 50s)
* 13:41 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:40 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1072
* 13:40 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:39 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-esams and A:durum
* 13:38 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1072
* 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:38 samtar@deploy1003: samtar, thiemowmde: Continuing with deployment
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-esams and A:durum
* 13:37 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:33 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:33 samtar@deploy1003: samtar, thiemowmde: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:32 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqsin and A:durum
* 13:31 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqsin and A:durum
* 13:28 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]]
* 13:28 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 13:27 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:24 ayounsi@cumin1004: START - Cookbook sre.hosts.reimage for host netflow2004.codfw.wmnet with OS trixie
* 13:24 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:22 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 13:22 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-esams
* 13:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:15 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:15 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 moritzm: installing bash updates from trixie point release
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 13:14 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1067
* 13:13 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1067
* 13:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2901: Test
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:09 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:09 moritzm: installing libxslt bugfix updates from Trixie point release
* 13:08 jelto@dns1004: END - running authdns-update
* 13:06 jelto@dns1004: START - running authdns-update
* 13:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 13:04 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:04 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 13:00 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:59 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1066
* 12:59 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1066
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:58 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:54 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 12:53 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2901: Test
* 12:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2901: Test
* 12:52 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'sync'.
* 12:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 12:50 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2901: Test
* 12:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:50 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 12:42 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:41 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-magru and A:durum
* 12:38 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-magru and A:durum
* 12:34 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 12:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1282: Pooling db1282 into s6
* 12:31 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 12:25 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:21 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:19 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:15 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_ulsfo
* 12:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1020.eqiad.wmnet with OS bookworm
* 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: db2207 repool
* 12:07 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_ulsfo
* 12:04 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 11:58 kart_: cxserver: Use urldownloader LVS endpoint ([[phab:T429175|T429175]])
* 11:57 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply
* 11:56 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply
* 11:56 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_magru
* 11:55 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply
* 11:55 moritzm: installing rsync security updates
* 11:55 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply
* 11:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:52 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply
* 11:51 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply
* 11:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1282: Pooling db1282 into s6
* 11:45 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1282 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96328 and previous config saved to /var/cache/conftool/dbconfig/20260903-114526-marostegui.json
* 11:43 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqiad
* 11:35 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqiad
* 11:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1020.eqiad.wmnet with OS bookworm
* 11:26 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 11:24 cgoubert@deploy1003: Finished scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter (duration: 12m 01s)
* 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: db2207 repool
* 11:22 cgoubert@deploy1003: cgoubert: Continuing with deployment
* 11:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqsin
* 11:17 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_esams
* 11:15 cgoubert@deploy1003: cgoubert: mediawiki: enable forward of fatal metrics to statsd exporter synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:14 cgoubert@deploy1003: Started scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter
* 11:10 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_esams
* 11:09 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 11:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1019.eqiad.wmnet with OS bookworm
* 11:01 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_drmrs
* 10:59 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_codfw
* 10:52 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_codfw
* 10:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:41 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 10:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 10:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow2005.codfw.wmnet
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2005.codfw.wmnet with OS trixie
* 10:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:17 btullis@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync
* 10:17 btullis@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync
* 10:16 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1019.eqiad.wmnet with OS bookworm
* 10:15 btullis@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: sync
* 10:15 btullis@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: sync
* 10:12 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s8
* 10:11 marostegui: Move s8 sanitarium from db1167 to db1281 [[phab:T434778|T434778]]
* 10:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 10:03 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 09:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:55 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2003.codfw.wmnet with OS trixie
* 09:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:43 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow2005.codfw.wmnet with OS trixie
* 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:42 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1018.eqiad.wmnet with OS bookworm
* 09:41 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_ulsfo
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:39 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2003.codfw.wmnet with reason: host reimage
* 09:39 hnowlan: fixed currently oncall pane in klaxon
* 09:38 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:38 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434751|T434751]]
* 09:37 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:37 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow2005.codfw.wmnet
* 09:35 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434775|T434775]]
* 09:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s7
* 09:34 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:33 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_ulsfo
* 09:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqiad
* 09:30 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] (duration: 09m 30s)
* 09:27 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:25 zabe@deploy1003: zabe: Continuing with deployment
* 09:25 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqiad
* 09:25 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 09:25 zabe@deploy1003: zabe: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1174 from dbctl [[phab:T436904|T436904]]', diff saved to https://phabricator.wikimedia.org/P96323 and previous config saved to /var/cache/conftool/dbconfig/20260903-092448-marostegui.json
* 09:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:21 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]]
* 09:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqsin
* 09:17 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:15 topranks: put traffic on Lumen codfw<->eqiad link as it is stable [[phab:T435810|T435810]]
* 09:14 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_esams
* 09:09 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:06 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_esams
* 09:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 09:03 marostegui: Move s6 sanitarium from db1165 to db1279 [[phab:T434775|T434775]]
* 09:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1018.eqiad.wmnet with OS bookworm
* 08:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s6
* 08:57 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_drmrs
* 08:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_codfw
* 08:55 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw
* 08:49 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_codfw
* 08:49 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:45 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 08:45 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:42 marostegui: Move s5 sanitarium from db1161 to db1275 [[phab:T434776|T434776]]
* 08:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 24 hosts with reason: Changing sanitarium master in s5
* 08:38 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_magru
* 08:37 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:37 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server (duration: 01m 21s)
* 08:37 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:36 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server
* 08:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:33 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server (duration: 01m 28s)
* 08:32 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server
* 08:27 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 08:11 moritzm: uploaded wmf-laptop 1.0.7 to apt.wikimedia.org
* 08:03 marostegui: Move s2 sanitarium from db1156 to db1271 [[phab:T434287|T434287]]
* 07:59 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2001.codfw.wmnet
* 07:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 23 hosts with reason: Changing sanitarium master in s2
* 07:49 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2001.codfw.wmnet
* 07:39 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:29 chlod: UTC morning backport window done
* 07:27 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] (duration: 11m 54s)
* 07:22 chlod@deploy1003: chlod, tryvix1509: Continuing with deployment
* 07:22 XioNoX: push pfw policies - [[phab:T436729|T436729]]
* 07:20 chlod@deploy1003: chlod, tryvix1509: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]]
* 07:15 marostegui: Power off db1228 for maintenance
* 07:13 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1228.eqiad.wmnet with reason: Onsite maintenance
* 07:01 arnaudb@dns1006: END - running authdns-update
* 06:58 arnaudb@dns1006: START - running authdns-update
* 06:54 jmm@cumin2003: END (PASS) - Cookbook sre.wdqs.restart-nginx-envoy (exit_code=0) rolling restart_daemons on A:wcqs-public
* 06:52 jmm@cumin2003: START - Cookbook sre.wdqs.restart-nginx-envoy rolling restart_daemons on A:wcqs-public
* 06:46 moritzm: installing libxml2 security updates
* 06:27 hashar: Upgrading CI Jenkins on contint1003 # [[phab:T436812|T436812]]
* 06:11 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet
* 06:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts2002.codfw.wmnet
* 06:00 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet
* 05:56 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts2002.codfw.wmnet
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 48s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:16 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1017.eqiad.wmnet with OS bookworm
== 2026-09-02 ==
* 23:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] (duration: 10m 21s)
* 23:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:50 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 23:49 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]]
* 23:38 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:38 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host aqs1017.eqiad.wmnet with OS bookworm
* 22:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] (duration: 11m 03s)
* 22:32 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 22:30 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:26 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]]
* 22:05 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] (duration: 14m 14s)
* 21:59 krinkle@deploy1003: krinkle: Continuing with deployment
* 21:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:50 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]]
* 21:44 inflatador: bking@apt1002 sudo -E private_reprepro --ignore=wrongdistribution -C matomo_plugins include bookworm-wikimedia-private matomo-plugin-customreports_5.5.0-1_amd64.changes [[phab:T431608|T431608]]
* 21:40 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] (duration: 09m 48s)
* 21:35 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:34 jforrester@deploy1003: jforrester: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:34 inflatador: bking@apt1002 sudo -E reprepro -C main include bookworm-wikimedia matomo-plugin-marketingcampaignsreporting_5.2.2-3_amd64.changes [[phab:T431608|T431608]]
* 21:30 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]]
* 21:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 21:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:03 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 20:50 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 20:27 dancy@deploy1003: Finished scap sync-world: testing (duration: 09m 21s)
* 20:18 dancy@deploy1003: Started scap sync-world: testing
* 20:18 dancy@deploy1003: Installation of scap version "4.288.0" completed for 3 hosts
* 20:16 dancy@deploy1003: Installing scap version "4.288.0" for 3 host(s)
* 19:57 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] (duration: 64m 27s)
* 19:55 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 18:57 jforrester@deploy1003: jforrester: Continuing with deployment
* 18:57 jforrester@deploy1003: jforrester: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:55 swfrench-wmf: deleted pods coredns-85b4f68d95-pk5sn coredns-85b4f68d95-22ddb coredns-85b4f68d95-49k5p in eqiad due to intermittent upstream resolution health check failures correlated with high DNS resolution latency
* 18:53 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]]
* 18:31 sukhe@dns1004: END - running authdns-update
* 18:28 sukhe@dns1004: START - running authdns-update
* 18:26 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:26 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:17 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:16 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 18:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:14 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:14 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqiad
* 18:14 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:49 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqiad
* 17:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:45 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqsin
* 17:38 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:38 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:20 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqsin
* 17:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 17:00 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 16:59 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:48 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 16:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:41 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 16:41 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 16:41 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:38 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:34 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 16:29 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_drmrs
* 16:24 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 16:14 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_drmrs
* 16:11 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-unlock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 16:10 root@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]] (duration: 48m 09s)
* 16:10 root@deploy1003: Forcefully removing global lock: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 16:10 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-unlock-scap for datacenter switchover from codfw to eqiad
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:59 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:58 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters for datacenter switchover from codfw to eqiad
* 15:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 15:58 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-restore-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-restore-ttl for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-start-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 root@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-start-maintenance for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-jobrunner: sync
* 15:56 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-jobrunner: sync
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.07-set-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period ends at: 2026-09-02 15:56:13.434320
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.07-set-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.03-set-db-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.03-set-db-readonly for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.02-set-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period starts at: 2026-09-02 15:53:47.690918
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.02-set-readonly for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.01-stop-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.01-stop-maintenance for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-reduce-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:47 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-reduce-ttl for datacenter switchover from codfw to eqiad
* 15:46 slyngshede@cumin1003: END (ERROR) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=97) for datacenter switchover from codfw to eqiad
* 15:45 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 15:44 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 15:42 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service
* 15:39 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:38 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:31 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:28 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:27 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 15:27 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from codfw to eqiad
* 15:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-lock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:22 root@deploy1003: Locking from deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-lock-scap for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks for datacenter switchover from codfw to eqiad
* 15:17 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service
* 15:15 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_magru
* 15:11 moritzm: import jenkins 2.568.3 to thirdparty/jenkins for trixie-wikimedia
* 14:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 14:44 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_magru
* 14:32 moritzm: installing pdns-recursor security updates
* 14:27 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:27 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:20 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:15 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 14:12 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 14:09 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:09 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:09 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:08 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:06 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:44 moritzm: bounce tcpircbot-logmsgbot/tcpircbot-logmsgbot_cloud on alert1002 to allow cumin1004 [[phab:T427897|T427897]]
* 13:36 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] (duration: 09m 52s)
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Continuing with deployment
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:26 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]]
* 13:25 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 13:24 moritzm: installing wireshark security updates
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 13:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:04 moritzm: import librsvg 2.60.0+dfsg-1+wmf13u1 to component/thumbor for trixie-wikimedia [[phab:T436505|T436505]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 12:44 atsuko@dns1004: END - running authdns-update
* 12:41 atsuko@dns1004: START - running authdns-update
* 12:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:35 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] (duration: 12m 50s)
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:24 dreamyjazz@deploy1003: dreamyjazz, btullis: Continuing with deployment
* 12:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:22 dreamyjazz@deploy1003: dreamyjazz, btullis: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:20 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:17 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]]
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 11:31 marostegui@cumin1003: Removing db1172 from zarcillo [[phab:T436763|T436763]]
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1172.eqiad.wmnet
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:30 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:26 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:26 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:25 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 11:25 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:24 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:20 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1172.eqiad.wmnet
* 11:20 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 11:12 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:11 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:10 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:09 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:05 slyngshede@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.*
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 11:05 slyngshede@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 11:03 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply
* 11:03 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
* 10:52 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:49 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:48 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Pool back db1242
* 10:45 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow4003.ulsfo.wmnet with OS trixie
* 10:31 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1172 from dbctl [[phab:T436763|T436763]]', diff saved to https://phabricator.wikimedia.org/P96318 and previous config saved to /var/cache/conftool/dbconfig/20260902-103152-marostegui.json
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 blake@deploy1003: Finished scap sync-world: non-build deployment for [[phab:T417800|T417800]] (duration: 05m 37s)
* 10:06 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:05 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:04 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:03 blake@deploy1003: Started scap sync-world: non-build deployment for [[phab:T417800|T417800]]
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 10:00 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:00 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:59 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:58 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:57 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:56 jmm@dns1004: END - running authdns-update
* 09:56 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:55 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:54 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:53 jmm@dns1004: START - running authdns-update
* 09:51 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:47 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1228 to dbctl [[phab:T435892|T435892]]', diff saved to https://phabricator.wikimedia.org/P96313 and previous config saved to /var/cache/conftool/dbconfig/20260902-094713-marostegui.json
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:42 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow4003.ulsfo.wmnet with OS trixie
* 09:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow7002.magru.wmnet with OS trixie
* 09:31 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:30 moritzm: installing openjdk-21 security updates
* 09:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:22 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:17 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:17 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:16 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:10 moritzm: installing openjdk-8 security updates
* 09:09 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:08 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:46 tappof: bump space for prometheus k8s-dse in eqiad
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:39 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow7002.magru.wmnet with OS trixie
* 08:36 moritzm: installing libgraphite2 security updates
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:23 Msz2001: UTC morning backport window done
* 08:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] (duration: 14m 36s)
* 08:19 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:19 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet
* 08:14 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:14 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet
* 08:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet
* 08:08 fabfur@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cp5022.eqsin.wmnet with reason: investigating
* 08:07 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]]
* 08:07 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp5022.*
* 08:07 fabfur: depooling and silencing cp5022 ([[phab:T414411|T414411]])
* 08:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet
* 08:03 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:03 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* {{safesubst:SAL entry|1=08:03 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)}}
* 07:49 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 07:49 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]] synced to the
* 07:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1004.eqiad.wmnet
* 07:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1004.eqiad.wmnet
* 07:30 jmm@dns1004: END - running authdns-update
* {{safesubst:SAL entry|1=07:27 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]}}
* 07:27 jmm@dns1004: START - running authdns-update
* 07:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] (duration: 16m 04s)
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1242: Cloning db1228
* 07:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1242: Cloning db1228
* 07:18 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1228,1242].eqiad.wmnet with reason: db1242 needs to clone db1228
* 07:17 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Continuing with deployment
* 07:12 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1228.eqiad.wmnet with OS trixie
* 07:10 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be veri
* 07:06 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]]
* 06:43 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: build: Updating npm dependencies (duration: 00m 13s)
* 06:43 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: build: Updating npm dependencies
* 06:39 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:32 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:23 slyngshede@dns1004: END - running authdns-update
* 06:21 marostegui: Drop cu* tables from s3 bswiktionary [[phab:T435965|T435965]]
* 06:20 slyngshede@dns1004: START - running authdns-update
* 06:18 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1228.eqiad.wmnet with OS trixie
* 06:13 XioNoX: re-enable magru cr1/asw1-b3 link - [[phab:T436675|T436675]]
* 05:06 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] (duration: 04m 42s)
* 05:04 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:03 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:01 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 05:01 tstarling@deploy1003: Scap cancelled without rolling back.
* 04:53 tstarling@deploy1003: tstarling: Continuing with deployment
* 04:29 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 04:25 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 43s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-01 ==
* 21:59 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] (duration: 18m 05s)
* 21:52 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:47 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]]
* 21:38 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] (duration: 23m 55s)
* 21:28 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:20 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:14 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]]
* 20:47 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1016.eqiad.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:24 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:11 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 20:11 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:57 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:57 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:56 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:54 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:52 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:47 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:45 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:42 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:41 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:40 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:39 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:32 jhancock@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudcephosd1055
* 19:32 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:30 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.* [reason: update IP addrs]
* 19:30 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 19:30 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:23 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:22 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:16 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:13 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1056
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1056
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1055
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:12 jclark@cumin1003: END (ERROR) - Cookbook sre.network.configure-switch-interfaces (exit_code=97) for host cloudcephosd1054
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1054
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:11 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:06 jclark@cumin1003: START - Cookbook sre.dns.netbox
* 19:06 sukhe@dns1004: END - running authdns-update
* 19:03 sukhe@dns1004: START - running authdns-update
* 18:14 dancy@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 17:04 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 17:00 dancy@deploy1003: Finished scap sync-world: testing (duration: 08m 07s)
* 16:52 dancy@deploy1003: Started scap sync-world: testing
* 16:48 dancy@deploy1003: sync-world aborted: testing (duration: 00m 05s)
* 16:48 dancy@deploy1003: Started scap sync-world: testing
* 16:47 dancy@deploy1003: Installation of scap version "4.287.0" completed for 156 hosts
* 16:42 dancy@deploy1003: Installing scap version "4.287.0" for 156 host(s)
* 16:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 16:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 15:51 moritzm: installing mesa security updates
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts phab1004.eqiad.wmnet
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:27 aokoth@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:20 aokoth@cumin1003: START - Cookbook sre.dns.netbox
* 15:14 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595] (duration: 05m 32s)
* 15:14 aokoth@cumin1003: START - Cookbook sre.hosts.decommission for hosts phab1004.eqiad.wmnet
* 15:11 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on phab1004.eqiad.wmnet with reason: Decom
* 15:09 joal@deploy1003: Started deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595]
* 15:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config: apply
* 15:08 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config: apply
* 14:55 hashar: Restarted Jenkins on releases1003
* 14:51 hashar: Restarted CI Jenkins on contint1003
* 14:48 hashar: Restarting Gerrit primary on gerrit2003
* 14:45 hashar: Restarted Gerrit on gerrit1003 and gerrit2002
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 14:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 14:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 14:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:24 moritzm: installing curl security updates
* 14:24 jmm@dns1004: END - running authdns-update
* 14:23 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser (duration: 00m 14s)
* 14:23 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:21 jmm@dns1004: START - running authdns-update
* 14:21 jmm@dns1004: END - running authdns-update
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 14:19 jmm@dns1004: START - running authdns-update
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet
* 14:14 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:14 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:13 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595] (duration: 07m 26s)
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Test
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Test
* 14:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet
* 14:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Test
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Test
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 14:05 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595]
* 14:05 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] (duration: 37m 37s)
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 hashar: Removed openjdk-17 packages from contint1002/contint2002 following relocation of CI Jenkins to contint1003/contint2003 # [[phab:T418521|T418521]]
* 14:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:02 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:02 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:00 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595] (duration: 00m 45s)
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 13:59 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595]
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:58 ladsgroup@dns1004: END - running authdns-update
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:57 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 ladsgroup@dns1004: START - running authdns-update
* 13:56 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:56 ladsgroup@dns1004: END - running authdns-update
* 13:55 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:53 ladsgroup@dns1004: START - running authdns-update
* 13:50 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:49 kharlan@deploy1003: kharlan: Continuing with deployment
* 13:48 kharlan@deploy1003: kharlan: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2004.wikimedia.org
* 13:42 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:41 jmm@dns1004: END - running authdns-update
* 13:40 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2004.wikimedia.org
* 13:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2003.wikimedia.org
* 13:38 jmm@dns1004: START - running authdns-update
* 13:34 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2003.wikimedia.org
* 13:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1004.wikimedia.org
* 13:26 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]]
* 13:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:25 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:24 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1004.wikimedia.org
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:23 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] (duration: 20m 24s)
* 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1003.wikimedia.org
* 13:20 fnegri@deploy1003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply
* 13:19 fnegri@deploy1003: helmfile [eqiad] START helmfile.d/services/toolhub: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:16 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1003.wikimedia.org
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:15 moritzm: bump urldownloader[12]00[34] to 8G RAM [[phab:T429175|T429175]]
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:15 fnegri@deploy1003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 fnegri@deploy1003: helmfile [codfw] START helmfile.d/services/toolhub: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:11 fnegri@deploy1003: helmfile [staging] DONE helmfile.d/services/toolhub: apply
* 13:11 aude@deploy1003: aude: Continuing with deployment
* 13:10 fnegri@deploy1003: helmfile [staging] START helmfile.d/services/toolhub: apply
* 13:07 aude@deploy1003: aude: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2004.wikimedia.org with OS bookworm
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow1004.eqiad.wmnet
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1004.eqiad.wmnet with OS trixie
* 12:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:47 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:41 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] (duration: 16m 25s)
* 12:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 12:34 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 12:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 12:33 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 12:31 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:29 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:25 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]]
* 12:22 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1004.eqiad.wmnet with OS trixie
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:20 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 12:20 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 12:16 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow1004.eqiad.wmnet
* 12:15 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2004.wikimedia.org with OS bookworm
* 12:14 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 12:07 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2003.wikimedia.org with OS bookworm
* 11:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:43 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:32 jmm@cumin2003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling restart_daemons on A:kafka-test-eqiad
* 11:26 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2003.wikimedia.org with OS bookworm
* 11:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1004.wikimedia.org with OS bookworm
* 11:12 moritzm: installing openjdk-21 security updates
* 11:12 jmm@cumin2003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling restart_daemons on A:kafka-test-eqiad
* 10:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:53 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:49 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2902: Pool back db2902
* 10:45 moritzm: installing Python 3.11 security updates
* 10:37 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1004.wikimedia.org with OS bookworm
* 10:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5022.eqsin.wmnet with OS trixie
* 10:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp5022.eqsin.wmnet on all recursors
* 10:36 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache cp5022.eqsin.wmnet on all recursors
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:30 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:04 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:04 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2902: Pool back db2902
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:03 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2902: test
* 10:03 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 10:01 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 09:58 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:50 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/echoserver: apply
* 09:49 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/echoserver: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Change db1176 and db2230's weight, test-s4 masters, to 0 to mimic the rest of production [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96292 and previous config saved to /var/cache/conftool/dbconfig/20260901-092444-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96291 and previous config saved to /var/cache/conftool/dbconfig/20260901-090233-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96290 and previous config saved to /var/cache/conftool/dbconfig/20260901-090158-marostegui.json
* 09:01 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96289 and previous config saved to /var/cache/conftool/dbconfig/20260901-090121-marostegui.json
* 09:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:56 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:50 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1003.wikimedia.org with OS bookworm
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1074.eqiad.wmnet
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:46 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:43 marostegui@cumin1003: dbctl commit (dc=all): 'Test repool db2902', diff saved to https://phabricator.wikimedia.org/P96288 and previous config saved to /var/cache/conftool/dbconfig/20260901-084317-marostegui.json
* 08:42 marostegui@cumin1003: dbctl commit (dc=all): 'Test depool db2902', diff saved to https://phabricator.wikimedia.org/P96287 and previous config saved to /var/cache/conftool/dbconfig/20260901-084249-marostegui.json
* 08:39 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:37 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2902: test
* 08:36 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96285 and previous config saved to /var/cache/conftool/dbconfig/20260901-083557-marostegui.json
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96284 and previous config saved to /var/cache/conftool/dbconfig/20260901-083527-marostegui.json
* 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96283 and previous config saved to /var/cache/conftool/dbconfig/20260901-083432-marostegui.json
* 08:32 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1074.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1073.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:27 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:24 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 08:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022']
* 08:16 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1073.eqiad.wmnet
* 08:15 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1003.wikimedia.org with OS bookworm
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1072.eqiad.wmnet
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:14 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:10 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:05 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1072.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1067.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:00 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:56 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:54 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022']
* 07:53 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 07:52 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie
* 07:52 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 07:50 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1067.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1066.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:47 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 07:41 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:34 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1066.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1065.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:32 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:27 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:23 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1065.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1075.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:15 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 06:49 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:45 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1075.eqiad.wmnet
* 06:29 moritzm: installing Java 17 security updates
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.15 (duration: 02m 25s)
* 03:50 denisse@deploy1003: Finished deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2 (duration: 00m 19s)
* 03:50 denisse@deploy1003: Started deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2
* 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]] (duration: 37m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:30 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 00:29 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
== Other archives ==
See [[Server Admin Log/Archives]].
<noinclude>
[[Category:SAL]]
[[Category:Operations]]
</noinclude>
nevqwbomuxugydkzsm43una876v1655
2458704
2458703
2026-09-19T16:55:34Z
Stashbot
7414
ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
2458704
wikitext
text/x-wiki
== 2026-09-19 ==
* 16:55 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:11 urbanecm: Attach SHB@commonswiki to the SUL account manually ([[phab:T438591|T438591]], see [[phab:T438591|T438591]]#12341750 for what I did exactly)
* 04:08 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 04:08 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 04:08 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 04:07 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 36s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-18 ==
* 22:41 rzl@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=sessionstore,name=eqiad
* 17:08 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343065{{!}}mw.DesktopArticleTarget: if source education is enabled suppress welcome (T434249)]] (duration: 09m 26s)
* 17:05 Dreamy_Jazz: Created `securepoll_log` on `nlwiki` main DB cluster for [[phab:T434045|T434045]]
* 17:04 kemayo@deploy1003: kemayo: Continuing with deployment
* 17:03 kemayo@deploy1003: kemayo: Backport for [[gerrit:1343065{{!}}mw.DesktopArticleTarget: if source education is enabled suppress welcome (T434249)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:59 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1343065{{!}}mw.DesktopArticleTarget: if source education is enabled suppress welcome (T434249)]]
* 16:49 oblivian@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343127{{!}}ResourceLoader: hotfix for current logspam over the weekend (T438387)]] (duration: 11m 36s)
* 16:42 oblivian@deploy1003: oblivian: Continuing with deployment
* 16:42 oblivian@deploy1003: oblivian: Backport for [[gerrit:1343127{{!}}ResourceLoader: hotfix for current logspam over the weekend (T438387)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:37 oblivian@deploy1003: Started scap sync-world: Backport for [[gerrit:1343127{{!}}ResourceLoader: hotfix for current logspam over the weekend (T438387)]]
* 16:09 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ml-serve1016.eqiad.wmnet with OS trixie
* 14:49 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie
* 13:37 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host registry1004.eqiad.wmnet with OS trixie
* 13:23 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on registry1004.eqiad.wmnet with reason: host reimage
* 13:18 elukey@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on registry1004.eqiad.wmnet with reason: host reimage
* 13:04 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host registry1004.eqiad.wmnet with OS trixie
* 12:14 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:13 jclark@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:30 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 09:30 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 09:22 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on an-redacteddb1001.eqiad.wmnet with reason: cloning
* 09:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 09:19 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 09:18 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 21 hosts with reason: cloning db1270
* 09:18 marostegui: clone db1270:x4 from db1155:x4 lag will appear on x4
* 09:09 brouberol@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 09:08 brouberol@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 08:23 brouberol@dns1004: END - running authdns-update
* 08:21 brouberol@dns1004: START - running authdns-update
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 57s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-17 ==
* 21:04 tsev@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T438395|T438395]]
* 20:55 tsev@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T438395|T438395]]
* 20:47 jhuneidi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342774{{!}}Worklist Promotion test kitchen - Enable flag in production (T434513)]], [[gerrit:1342781{{!}}Exclude returntoapp query from app interception on iOS (T438395)]], [[gerrit:1342798{{!}}Revert "Update wikimania wordmark for 2026"]] (duration: 35m 59s)
* 20:35 jhuneidi@deploy1003: robertsky, jhuneidi, cmelo, tsev: Continuing with deployment
* 20:31 jhuneidi@deploy1003: robertsky, jhuneidi, cmelo, tsev: Backport for [[gerrit:1342774{{!}}Worklist Promotion test kitchen - Enable flag in production (T434513)]], [[gerrit:1342781{{!}}Exclude returntoapp query from app interception on iOS (T438395)]], [[gerrit:1342798{{!}}Revert "Update wikimania wordmark for 2026"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:11 jhuneidi@deploy1003: Started scap sync-world: Backport for [[gerrit:1342774{{!}}Worklist Promotion test kitchen - Enable flag in production (T434513)]], [[gerrit:1342781{{!}}Exclude returntoapp query from app interception on iOS (T438395)]], [[gerrit:1342798{{!}}Revert "Update wikimania wordmark for 2026"]]
* 19:15 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1100.eqiad.wmnet with OS trixie
* 19:15 vriley@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1004"
* 19:10 vriley@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1004"
* 18:52 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1100.eqiad.wmnet with reason: host reimage
* 18:48 vriley@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1100.eqiad.wmnet with reason: host reimage
* 18:32 vriley@cumin1004: START - Cookbook sre.hosts.reimage for host ms-be1100.eqiad.wmnet with OS trixie
* 18:20 urbanecm: Deploy a security fix for [[phab:T438389|T438389]]
* 17:47 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be1100.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:38 vriley@cumin1004: START - Cookbook sre.hosts.provision for host ms-be1100.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:38 vriley@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1100.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:37 vriley@cumin1004: START - Cookbook sre.hosts.provision for host ms-be1100.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:23 vriley@cumin1004: START - Cookbook sre.hosts.reimage for host ms-be1100.eqiad.wmnet with OS trixie
* 16:52 aokoth@deploy1003: Finished deploy [phabricator/deployment@c386249]: Deploy Phab (duration: 00m 12s)
* 16:52 aokoth@deploy1003: Started deploy [phabricator/deployment@c386249]: Deploy Phab
* 16:42 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1099.eqiad.wmnet with OS trixie
* 16:42 vriley@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1004"
* 16:42 vriley@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1004"
* 16:35 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be1100.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:31 aokoth@deploy1003: Finished deploy [phabricator/deployment@c386249]: Deploy Phab (duration: 00m 19s)
* 16:31 aokoth@deploy1003: Started deploy [phabricator/deployment@c386249]: Deploy Phab
* 16:21 vriley@cumin1004: START - Cookbook sre.hosts.provision for host ms-be1100.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:20 vriley@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be1100
* 16:20 vriley@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host ms-be1100
* 16:19 vriley@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:19 vriley@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [ms-be1100] - vriley@cumin1004"
* 16:19 vriley@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [ms-be1100] - vriley@cumin1004"
* 16:15 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 16:15 vriley@cumin1004: START - Cookbook sre.dns.netbox
* 16:15 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 16:14 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 16:14 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 16:12 moritzm: installing libapache-mod-auth-oidc security updates
* 16:12 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1099.eqiad.wmnet with reason: host reimage
* 16:11 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 16:11 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 16:08 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/editcheck-headless: apply
* 16:07 rzl@deploy1003: helmfile [staging] START helmfile.d/services/editcheck-headless: apply
* 16:06 vriley@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1099.eqiad.wmnet with reason: host reimage
* 16:01 moritzm: installing aom security updates
* 16:01 btullis@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ceph-admin2001.codfw.wmnet
* 16:01 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ceph-admin2001.codfw.wmnet with OS bookworm
* 15:55 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 15:55 rzl@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 15:54 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 15:54 rzl@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 15:54 rzl@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 15:53 rzl@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 15:53 rzl@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 15:52 rzl@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 15:51 vriley@cumin1004: START - Cookbook sre.hosts.reimage for host ms-be1099.eqiad.wmnet with OS trixie
* 15:48 moritzm: installing libde265 security updates
* 15:44 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ceph-admin2001.codfw.wmnet with reason: host reimage
* 15:40 btullis@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ceph-admin1001.eqiad.wmnet
* 15:40 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ceph-admin1001.eqiad.wmnet with OS bookworm
* 15:39 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir4003.*
* 15:37 btullis@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ceph-admin2001.codfw.wmnet with reason: host reimage
* 15:26 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir4003.ulsfo.wmnet with OS trixie
* 15:25 aokoth@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host phab2003.codfw.wmnet with OS trixie
* 15:23 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ceph-admin1001.eqiad.wmnet with reason: host reimage
* 15:19 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host registry1005.eqiad.wmnet with OS trixie
* 15:17 btullis@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ceph-admin1001.eqiad.wmnet with reason: host reimage
* 15:16 btullis@cumin1004: START - Cookbook sre.hosts.reimage for host ceph-admin2001.codfw.wmnet with OS bookworm
* 15:16 btullis@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ceph-admin2001.codfw.wmnet - btullis@cumin1004"
* 15:16 btullis@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ceph-admin2001.codfw.wmnet - btullis@cumin1004"
* 15:15 btullis@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ceph-admin2001.codfw.wmnet on all recursors
* 15:15 btullis@cumin1004: START - Cookbook sre.dns.wipe-cache ceph-admin2001.codfw.wmnet on all recursors
* 15:15 btullis@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:15 btullis@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ceph-admin2001.codfw.wmnet - btullis@cumin1004"
* 15:15 btullis@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ceph-admin2001.codfw.wmnet - btullis@cumin1004"
* 15:08 aokoth@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on phab2003.codfw.wmnet with reason: host reimage
* 15:06 Msz2001: Deployed private code changes to Suggestedinvestigations
* 15:05 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir4003.ulsfo.wmnet with reason: host reimage
* 15:05 aokoth@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on phab2003.codfw.wmnet with reason: host reimage
* 15:04 btullis@cumin1004: START - Cookbook sre.hosts.reimage for host ceph-admin1001.eqiad.wmnet with OS bookworm
* 15:04 btullis@cumin1004: START - Cookbook sre.dns.netbox
* 15:04 btullis@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ceph-admin1001.eqiad.wmnet - btullis@cumin1004"
* 15:04 btullis@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ceph-admin1001.eqiad.wmnet - btullis@cumin1004"
* 15:04 btullis@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ceph-admin1001.eqiad.wmnet on all recursors
* 15:04 btullis@cumin1004: START - Cookbook sre.dns.wipe-cache ceph-admin1001.eqiad.wmnet on all recursors
* 15:04 btullis@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:04 btullis@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ceph-admin1001.eqiad.wmnet - btullis@cumin1004"
* 15:04 btullis@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ceph-admin1001.eqiad.wmnet - btullis@cumin1004"
* 15:02 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on registry1005.eqiad.wmnet with reason: host reimage
* 15:01 btullis@cumin1004: START - Cookbook sre.ganeti.makevm for new host ceph-admin2001.codfw.wmnet
* 15:00 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342265{{!}}JsonSchemaBuilder: Cache the root schema in the process (T437588)]], [[gerrit:1342264{{!}}JsonSchemaBuilder: Cache the root schema in the process (T437588)]], [[gerrit:1342684{{!}}SI: Preserve the username filter when switching queues (T438308)]] (duration: 12m 34s)
* 15:00 btullis@cumin1004: START - Cookbook sre.dns.netbox
* 15:00 btullis@cumin1004: START - Cookbook sre.ganeti.makevm for new host ceph-admin1001.eqiad.wmnet
* 14:58 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir4003.ulsfo.wmnet with reason: host reimage
* 14:58 elukey@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on registry1005.eqiad.wmnet with reason: host reimage
* 14:55 urbanecm@deploy1003: mszwarc, urbanecm: Continuing with deployment
* 14:51 urbanecm@deploy1003: mszwarc, urbanecm: Backport for [[gerrit:1342265{{!}}JsonSchemaBuilder: Cache the root schema in the process (T437588)]], [[gerrit:1342264{{!}}JsonSchemaBuilder: Cache the root schema in the process (T437588)]], [[gerrit:1342684{{!}}SI: Preserve the username filter when switching queues (T438308)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:51 aokoth@cumin1004: START - Cookbook sre.hosts.reimage for host phab2003.codfw.wmnet with OS trixie
* 14:50 aokoth@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on phab2003.codfw.wmnet with reason: Reimage
* 14:47 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1342265{{!}}JsonSchemaBuilder: Cache the root schema in the process (T437588)]], [[gerrit:1342264{{!}}JsonSchemaBuilder: Cache the root schema in the process (T437588)]], [[gerrit:1342684{{!}}SI: Preserve the username filter when switching queues (T438308)]]
* 14:42 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1331830{{!}}Keep Balinese Palm Leaf variants enabled on wikisource (T436398)]], [[gerrit:1340216{{!}}Turn on variant conversion for PageAssessments (T328012)]], [[gerrit:1341949{{!}}Parsoid Read Views: Enable on 61 wikiquote wikis (T437917)]] (duration: 15m 31s)
* 14:39 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host registry1005.eqiad.wmnet with OS trixie
* 14:35 cscott@deploy1003: ssastry, cscott: Continuing with deployment
* 14:33 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir4003.ulsfo.wmnet with OS trixie
* 14:32 cscott@deploy1003: ssastry, cscott: Backport for [[gerrit:1331830{{!}}Keep Balinese Palm Leaf variants enabled on wikisource (T436398)]], [[gerrit:1340216{{!}}Turn on variant conversion for PageAssessments (T328012)]], [[gerrit:1341949{{!}}Parsoid Read Views: Enable on 61 wikiquote wikis (T437917)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:26 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1331830{{!}}Keep Balinese Palm Leaf variants enabled on wikisource (T436398)]], [[gerrit:1340216{{!}}Turn on variant conversion for PageAssessments (T328012)]], [[gerrit:1341949{{!}}Parsoid Read Views: Enable on 61 wikiquote wikis (T437917)]]
* 14:20 caro@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342678{{!}}enwiki desktop VE: add education popup for switching to source editor (T434249)]], [[gerrit:1342362{{!}}Make VE the default editor on enwiki desktop (T436574)]] (duration: 33m 52s)
* 14:07 caro@deploy1003: caro: Continuing with deployment
* 14:06 caro@deploy1003: caro: Backport for [[gerrit:1342678{{!}}enwiki desktop VE: add education popup for switching to source editor (T434249)]], [[gerrit:1342362{{!}}Make VE the default editor on enwiki desktop (T436574)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:46 caro@deploy1003: Started scap sync-world: Backport for [[gerrit:1342678{{!}}enwiki desktop VE: add education popup for switching to source editor (T434249)]], [[gerrit:1342362{{!}}Make VE the default editor on enwiki desktop (T436574)]]
* 13:37 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 13:34 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:33 elukey@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:31 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:30 elukey@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:29 Emperor: apus - radosgw-admin quota set --quota-scope=user --uid=docker-registry --max-size=5T [[phab:T438339|T438339]]
* 13:24 jclark@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1016
* 13:24 jclark@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1016
* 13:11 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:11 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:09 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:09 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:55 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:53 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 12:42 jclark@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:40 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:40 jclark@cumin1003: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:35 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 12:19 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342626{{!}}MathMathML: Simplify Mathoid fallback/a11y class logic (T436026)]], [[gerrit:1342627{{!}}ext.math.mathjax: Implement mwe-math-mathml-a11y for client-side MathJax (T436026)]] (duration: 13m 04s)
* 12:14 krinkle@deploy1003: krinkle: Continuing with deployment
* 12:10 krinkle@deploy1003: krinkle: Backport for [[gerrit:1342626{{!}}MathMathML: Simplify Mathoid fallback/a11y class logic (T436026)]], [[gerrit:1342627{{!}}ext.math.mathjax: Implement mwe-math-mathml-a11y for client-side MathJax (T436026)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:05 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1342626{{!}}MathMathML: Simplify Mathoid fallback/a11y class logic (T436026)]], [[gerrit:1342627{{!}}ext.math.mathjax: Implement mwe-math-mathml-a11y for client-side MathJax (T436026)]]
* 10:37 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 10:28 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 10:10 blake@deploy1003: Finished scap sync-world: cleanup for [[phab:T417800|T417800]] (duration: 03m 57s)
* 10:07 blake@deploy1003: Started scap sync-world: cleanup for [[phab:T417800|T417800]]
* 09:52 marostegui@cumin1004: dbctl commit (dc=all): 'Fix weights [[phab:T436496|T436496]]', diff saved to https://phabricator.wikimedia.org/P96466 and previous config saved to /var/cache/conftool/dbconfig/20260917-095235-marostegui.json
* 09:51 marostegui@cumin1004: dbctl commit (dc=all): 'Fix weights [[phab:T436496|T436496]]', diff saved to https://phabricator.wikimedia.org/P96465 and previous config saved to /var/cache/conftool/dbconfig/20260917-095131-marostegui.json
* 09:41 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:41 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 09:41 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:40 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 09:40 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:40 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 09:35 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 09:35 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 08:37 moritzm: pruned obsolete Bullseye image prometheus-nutcracker-exporter from the docker registry [[phab:T416452|T416452]]
* 08:34 XioNoX: Manually install gnmic 0.49.0 on netflow2005 - [[phab:T438291|T438291]]
* 08:28 brouberol@dns1004: END - running authdns-update
* 08:26 brouberol@dns1004: START - running authdns-update
* 08:13 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.20 refs [[phab:T430839|T430839]]
* 08:10 moritzm: imported nodejs_26.8.2-1nodesource1 to thirdparty/node26 for trixie-wikimedia [[phab:T437510|T437510]]
* 08:07 Amir1: dropped links tables from db2206 ([[phab:T437278|T437278]])
* 08:03 Amir1: dropped links tables from db2219 ([[phab:T437278|T437278]])
* 08:01 Amir1: dropped links tables from db2236 ([[phab:T437278|T437278]])
* 07:59 Amir1: dropped non-links tables from db1262 ([[phab:T437278|T437278]])
* 07:57 Amir1: dropped non-links tables from db2245 ([[phab:T437278|T437278]])
* 07:52 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@8eaca67] (releasing): [[phab:T438205|T438205]] to prod host (duration: 00m 44s)
* 07:52 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@8eaca67] (releasing): [[phab:T438205|T438205]] to prod host
* 07:49 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@8eaca67] (releasing): [[phab:T438205|T438205]] to backup host (duration: 00m 47s)
* 07:48 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@8eaca67] (releasing): [[phab:T438205|T438205]] to backup host
* 07:25 XioNoX: Manually install gnmic 0.49.0 on netflow1004 - [[phab:T438291|T438291]]
* 07:24 mlitn@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342398{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342396{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342395{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342545{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342546{{!}}Localisation updates from https://translatewiki.net.]],
* 07:19 mlitn@deploy1003: mlitn, jdlrobson: Continuing with deployment
* {{safesubst:SAL entry|1=07:18 mlitn@deploy1003: mlitn, jdlrobson: Backport for [[gerrit:1342398{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342396{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342395{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342545{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342546{{!}}Localisation updates from https://translatewiki.net.]], [[gerri}}
* 07:11 mlitn@deploy1003: Started scap sync-world: Backport for [[gerrit:1342398{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342396{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342395{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342545{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342546{{!}}Localisation updates from https://translatewiki.net.]],
* 07:10 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Jmoore111 out of all services on: 2444 hosts
* 06:07 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 05:53 marostegui@cumin1004: END (FAIL) - Cookbook sre.mysql.decommission (exit_code=99)
* 05:53 marostegui@cumin1004: Removing db1180 from zarcillo [[phab:T437222|T437222]]
* 05:53 marostegui@cumin1004: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1180.eqiad.wmnet
* 05:53 marostegui@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 05:53 marostegui@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1180.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1004"
* 05:53 marostegui@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1180.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1004"
* 05:49 marostegui@cumin1004: START - Cookbook sre.dns.netbox
* 05:44 marostegui@cumin1004: START - Cookbook sre.hosts.decommission for hosts db1180.eqiad.wmnet
* 05:43 marostegui@cumin1004: START - Cookbook sre.mysql.decommission
* 04:26 aokoth@cumin1004: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet
* 04:24 aokoth@cumin1004: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet
== 2026-09-16 ==
* 23:10 rzl: rzl@deploy1003 Finished scap sync-world: Backport for [[gerrit:1342091{{!}}Repool poolcounter[1007,2006] (T435163)]] (duration: 11m 09s)
* 22:50 rzl@deploy1003: Started scap sync-world: Backport for [[gerrit:1342091{{!}}Repool poolcounter[1007,2006] (T435163)]]
* 22:47 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342383{{!}}DonorIdentification: Confirm before unlinking donor status in preferences (T436698)]], [[gerrit:1342385{{!}}Make learn more link to new window (T438252)]] (duration: 35m 21s)
* 22:35 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 22:33 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1342383{{!}}DonorIdentification: Confirm before unlinking donor status in preferences (T436698)]], [[gerrit:1342385{{!}}Make learn more link to new window (T438252)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:12 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1342383{{!}}DonorIdentification: Confirm before unlinking donor status in preferences (T436698)]], [[gerrit:1342385{{!}}Make learn more link to new window (T438252)]]
* 22:10 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2006.codfw.wmnet
* 22:06 rzl@cumin2003: START - Cookbook sre.hosts.reboot-single for host poolcounter2006.codfw.wmnet
* 22:06 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1007.eqiad.wmnet
* 22:02 rzl@cumin2003: START - Cookbook sre.hosts.reboot-single for host poolcounter1007.eqiad.wmnet
* 21:56 rzl@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342090{{!}}Repool poolcounter[1006,2005]; depool poolcounter[1007,2006] for reboot (T435163)]] (duration: 09m 39s)
* 21:52 rzl@deploy1003: rzl: Continuing with deployment
* 21:51 rzl@deploy1003: rzl: Backport for [[gerrit:1342090{{!}}Repool poolcounter[1006,2005]; depool poolcounter[1007,2006] for reboot (T435163)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:47 rzl@deploy1003: Started scap sync-world: Backport for [[gerrit:1342090{{!}}Repool poolcounter[1006,2005]; depool poolcounter[1007,2006] for reboot (T435163)]]
* 21:46 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:43 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2005.codfw.wmnet
* 21:42 vriley@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1099.eqiad.wmnet with OS trixie
* 21:41 vriley@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1098.eqiad.wmnet with OS trixie
* 21:41 vriley@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin2003"
* 21:40 vriley@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin2003"
* 21:40 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:40 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:39 rzl@cumin2003: START - Cookbook sre.hosts.reboot-single for host poolcounter2005.codfw.wmnet
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1006.eqiad.wmnet
* 21:38 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:38 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:37 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:35 rzl@cumin2003: START - Cookbook sre.hosts.reboot-single for host poolcounter1006.eqiad.wmnet
* 21:32 rzl@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342089{{!}}Depool poolcounter[1006,2005] for reboot (T435163)]] (duration: 13m 53s)
* 21:26 rzl@deploy1003: rzl: Continuing with deployment
* 21:25 rzl@deploy1003: rzl: Backport for [[gerrit:1342089{{!}}Depool poolcounter[1006,2005] for reboot (T435163)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:23 vriley@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1098.eqiad.wmnet with reason: host reimage
* 21:18 rzl@deploy1003: Started scap sync-world: Backport for [[gerrit:1342089{{!}}Depool poolcounter[1006,2005] for reboot (T435163)]]
* 21:17 vriley@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1098.eqiad.wmnet with reason: host reimage
* 21:10 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:09 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:09 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:09 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:08 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:02 vriley@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS trixie
* 20:49 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342339{{!}}Reapply "Tell VisualEditor about the app web edit tags", modified]] (duration: 35m 55s)
* 20:37 kemayo@deploy1003: cklimas, kemayo: Continuing with deployment
* 20:33 kemayo@deploy1003: cklimas, kemayo: Backport for [[gerrit:1342339{{!}}Reapply "Tell VisualEditor about the app web edit tags", modified]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:13 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1342339{{!}}Reapply "Tell VisualEditor about the app web edit tags", modified]]
* 19:20 dwisehaupt@dns1005: END - running authdns-update
* 19:18 dwisehaupt@dns1005: START - running authdns-update
* 19:06 vriley@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:57 dwisehaupt@dns1005: END - running authdns-update
* 18:55 vriley@cumin2003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:55 dwisehaupt@dns1005: START - running authdns-update
* 18:44 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 18:42 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 18:37 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 18:35 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 18:26 robh@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:22 robh@cumin2003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:21 dzahn@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 18:21 robh@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:21 robh@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:20 dzahn@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 18:20 dzahn@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 18:18 dzahn@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 18:18 mutante: k8s/miscweb: admin_ng deploy: creating namespace for attribution.wikimedia.org [[phab:T437635|T437635]]
* 18:17 dzahn@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 18:17 dzahn@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 18:17 dzahn@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 18:16 dzahn@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 18:13 cdanis@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "fix known-client creation - cdanis@cumin1003"
* 18:13 cdanis@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: fix known-client creation - cdanis@cumin1003
* 18:12 cdanis@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: fix known-client creation - cdanis@cumin1003
* 18:12 cdanis@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "fix known-client creation - cdanis@cumin1003"
* 18:04 vriley@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:51 vriley@cumin2003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:46 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 17:46 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 17:45 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:45 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:44 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:44 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:35 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1006.eqiad.wmnet with OS trixie
* 17:35 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1003"
* 17:30 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:29 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1003"
* 17:27 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:27 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:23 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:23 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:20 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 17:20 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 17:14 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:13 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1006.eqiad.wmnet with reason: host reimage
* 17:10 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:10 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:10 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:10 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:09 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:07 jclark@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1006.eqiad.wmnet with reason: host reimage
* 17:07 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 17:05 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:05 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:52 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie
* 16:46 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir7003.magru.wmnet
* 16:44 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir7003
* 16:19 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie
* 15:42 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply
* 15:42 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply
* 15:42 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 15:42 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir7003.magru.wmnet with OS trixie
* 15:41 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 15:36 jnuche@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342279{{!}}Use parser output value instead of status (T438154)]] (duration: 33m 21s)
* 15:35 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply
* 15:35 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply
* 15:24 jnuche@deploy1003: jnuche, jforrester: Continuing with deployment
* 15:23 jnuche@deploy1003: jnuche, jforrester: Backport for [[gerrit:1342279{{!}}Use parser output value instead of status (T438154)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:19 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie
* 15:18 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir7003.magru.wmnet with reason: host reimage
* 15:14 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir7003.magru.wmnet with reason: host reimage
* 15:03 jnuche@deploy1003: Started scap sync-world: Backport for [[gerrit:1342279{{!}}Use parser output value instead of status (T438154)]]
* 14:50 moritzm: installing apache2 security updates
* 14:49 slyngshede@cumin1003: conftool action : set/pooled=yes; selector: name=cp5026.eqsin.wmnet
* 14:47 slyngshede@cumin1003: conftool action : set/weight=1; selector: name=cp5026.eqsin.wmnet
* 14:45 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5026.eqsin.wmnet with OS trixie
* 14:44 moritzm: installing python-filelock security updates
* 14:42 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir7003.magru.wmnet with OS trixie
* 14:35 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:35 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:34 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:33 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp6002.drmrs.wmnet
* 14:32 sukhe@puppetserver1001: conftool action : set/weight=100; selector: name=cp6002.drmrs.wmnet,service=ats-be
* 14:32 sukhe@puppetserver1001: conftool action : set/weight=1; selector: name=cp6002.drmrs.wmnet,service=cdn
* 14:29 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6002.drmrs.wmnet with OS trixie
* 14:24 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: sync
* 14:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:24 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:24 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: sync
* 14:14 jforrester@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.19,1.47.0-wmf.20,next --multiversion-image-basename docker-registry.discovery.wmnet/restricte
* 14:14 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:14 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:13 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1342008{{!}}abstractwiki: Add three new articles per community advice to show off the feature (T434227)]]
* 14:10 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: sync
* 14:10 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: sync
* 14:10 jforrester@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.19,1.47.0-wmf.20,next --multiversion-image-basename docker-registry.discovery.wmnet/restricte
* 14:10 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: sync
* 14:10 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: sync
* 14:09 Amir1: dropped links tables on db2237 ([[phab:T437278|T437278]])
* 14:08 Amir1: dropped links tables on db1238 ([[phab:T437278|T437278]])
* 14:07 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1342008{{!}}abstractwiki: Add three new articles per community advice to show off the feature (T434227)]]
* 14:03 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5026.eqsin.wmnet with reason: host reimage
* 14:02 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:02 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:02 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 14:01 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:00 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 13:59 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6002.drmrs.wmnet with reason: host reimage
* 13:56 slyngshede@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5026.eqsin.wmnet with reason: host reimage
* 13:54 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cp6002.drmrs.wmnet with reason: host reimage
* 13:53 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 13:52 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 13:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dse-k8s-etcd[1001-1003].eqiad.wmnet
* 13:48 bking@cumin2003: START - Cookbook sre.hosts.remove-downtime for dse-k8s-etcd[1001-1003].eqiad.wmnet
* 13:46 bking@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM dse-k8s-etcd1001.eqiad.wmnet
* 13:46 bking@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM dse-k8s-etcd1001.eqiad.wmnet
* 13:45 bking@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM dse-k8s-etcd1002.eqiad.wmnet
* 13:41 bking@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM dse-k8s-etcd1002.eqiad.wmnet
* 13:41 bking@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM dse-k8s-etcd1003.eqiad.wmnet
* 13:38 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp6002.drmrs.wmnet with OS trixie
* 13:37 bking@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM dse-k8s-etcd1003.eqiad.wmnet
* 13:37 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.reboot-vm (exit_code=99) for VM dse-k8s-etcd1003.eqiad.wmnet
* 13:37 bking@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM dse-k8s-etcd1003.eqiad.wmnet
* 13:34 slyngshede@cumin1003: START - Cookbook sre.hosts.reimage for host cp5026.eqsin.wmnet with OS trixie
* 13:34 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 13:33 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5026.mgmt.eqsin.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:29 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp6002.mgmt.drmrs.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:25 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on dse-k8s-etcd[1001-1003].eqiad.wmnet with reason: Maintenance to increase vCPUS [[phab:T438084|T438084]]
* 13:24 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:24 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:22 slyngshede@cumin1003: START - Cookbook sre.hosts.provision for host cp5026.mgmt.eqsin.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342231{{!}}SI: Unset all filters on links to cases (T434530)]], [[gerrit:1342234{{!}}SI: Unset all filters on links to cases (T434530)]] (duration: 13m 10s)
* 13:19 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org
* 13:19 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org
* 13:19 sukhe@cumin1004: START - Cookbook sre.hosts.provision for host cp6002.mgmt.drmrs.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:18 stran@deploy1003: stran: Continuing with deployment
* 13:15 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: apply
* 13:15 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: apply
* 13:13 stran@deploy1003: stran: Backport for [[gerrit:1342231{{!}}SI: Unset all filters on links to cases (T434530)]], [[gerrit:1342234{{!}}SI: Unset all filters on links to cases (T434530)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:11 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp5026.eqsin.wmnet
* 13:09 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1342231{{!}}SI: Unset all filters on links to cases (T434530)]], [[gerrit:1342234{{!}}SI: Unset all filters on links to cases (T434530)]]
* 13:09 sukhe@cumin1004: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp6002.drmrs.wmnet
* 13:05 sukhe@cumin1004: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp6002.drmrs.wmnet
* 13:05 sukhe@cumin1004: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts cp6002.drmrs.wmnet
* 13:00 dkertesz@cumin1004: conftool action : set/pooled=yes; selector: name=cp5025.eqsin.wmnet
* 12:59 dkertesz@cumin1004: conftool action : set/weight=1; selector: name=cp5025.eqsin.wmnet
* 12:52 sukhe@cumin1004: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp6002.drmrs.wmnet
* 12:52 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cp6002.mgmt.drmrs.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:51 dkertesz: eqsin pooled again ([[phab:T438052|T438052]])
* 12:49 dkertesz@cumin1004: conftool action : set/pooled=yes; selector: cluster=dnsbox,dc=eqsin
* 12:47 dkertesz@dns1004: END - running authdns-update
* 12:45 dkertesz@dns1004: START - running authdns-update
* 12:43 dkertesz@cumin1004: conftool action : set/pooled=yes; selector: cluster=dnsbox,dc=eqsin,service=authdns-update
* 12:41 dkertesz@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqsin [reason: no reason specified, [[phab:T438052|T438052]]]
* 12:41 dkertesz@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool eqsin [reason: no reason specified, [[phab:T438052|T438052]]]
* 12:38 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a1-eqiad
* 12:38 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a1-eqiad
* 12:34 sukhe@cumin1004: START - Cookbook sre.hosts.provision for host cp6002.mgmt.drmrs.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:34 sukhe@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cp6002.drmrs.wmnet with reason: reimage
* 12:33 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp6002.drmrs.wmnet
* 12:13 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5025.eqsin.wmnet with OS trixie
* 12:12 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 12:11 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-a1-eqiad
* 12:09 cmooney@cumin1004: START - Cookbook sre.network.tls for network device ssw1-a1-eqiad
* 12:01 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.20 refs [[phab:T430839|T430839]]
* 11:59 moritzm: pruned obsolete Bullseye image python3-bullseye from the docker registry [[phab:T416452|T416452]]
* 11:50 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341284{{!}}IS/IS-labs: Set wmgUseModeratorToolkit default false (T431000)]] (duration: 10m 32s)
* 11:46 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ml-lab1002.eqiad.wmnet
* 11:45 samtar@deploy1003: samtar: Continuing with deployment
* 11:44 samtar@deploy1003: samtar: Backport for [[gerrit:1341284{{!}}IS/IS-labs: Set wmgUseModeratorToolkit default false (T431000)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:41 klausman@cumin1003: START - Cookbook sre.hosts.reboot-single for host ml-lab1002.eqiad.wmnet
* 11:39 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1341284{{!}}IS/IS-labs: Set wmgUseModeratorToolkit default false (T431000)]]
* 11:39 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5025.eqsin.wmnet with reason: host reimage
* 11:35 slyngshede@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5025.eqsin.wmnet with reason: host reimage
* 11:34 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:33 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:31 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:31 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:27 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:27 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:26 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:25 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:24 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.20 refs [[phab:T430839|T430839]]
* 11:21 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply
* 11:21 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply
* 11:20 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply
* 11:20 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply
* 11:19 moritzm: kicked off a new run of production-images-weekly-rebuild.service on build2004 (previously some leftovers of buster in the config prevented a complete run)
* 11:17 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
* 11:16 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
* 11:11 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
* 11:11 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
* 11:10 jnuche@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342210{{!}}Revert "Tell VisualEditor about the app web edit tags" (T437736 T438125)]] (duration: 33m 14s)
* 11:10 slyngshede@cumin1003: START - Cookbook sre.hosts.reimage for host cp5025.eqsin.wmnet with OS trixie
* 11:05 marostegui@cumin1004: dbctl commit (dc=all): 'Remove db1180 from dbctl [[phab:T437222|T437222]]', diff saved to https://phabricator.wikimedia.org/P96459 and previous config saved to /var/cache/conftool/dbconfig/20260916-110502-marostegui.json
* 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 11:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:01 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 11:01 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 10:57 jnuche@deploy1003: jnuche: Continuing with deployment
* 10:57 jnuche@deploy1003: jnuche: Backport for [[gerrit:1342210{{!}}Revert "Tell VisualEditor about the app web edit tags" (T437736 T438125)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:37 jnuche@deploy1003: Started scap sync-world: Backport for [[gerrit:1342210{{!}}Revert "Tell VisualEditor about the app web edit tags" (T437736 T438125)]]
* 10:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5025.mgmt.eqsin.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:10 slyngshede@cumin1003: START - Cookbook sre.hosts.provision for host cp5025.mgmt.eqsin.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:02 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 10:02 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 09:58 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 09:49 slyngshede@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on cp5025.eqsin.wmnet with reason: reimaging
* 09:48 slyngshede@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cp5025.eqsin.wmnet with reason: reimaging
* 09:41 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-timeline: apply
* 09:41 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-timeline: apply
* 09:38 moritzm: imported routinator 0.15.2-1trixie to thirdparty/routinator [[phab:T438122|T438122]]
* 09:30 slyngshede@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cp5025.mgmt.eqsin.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:29 slyngshede@cumin1003: START - Cookbook sre.hosts.provision for host cp5025.mgmt.eqsin.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:19 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 09:19 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp5025.eqsin.wmnet
* 09:19 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'sync'.
* 09:18 slyngshede@cumin1003: conftool action : set/pooled=yes; selector: name=cp3074.esams.wmnet
* 09:18 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp3074.esams.wmnet
* 09:15 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 09:12 elukey@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'sync'.
* 09:12 elukey@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'sync'.
* 09:11 elukey@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'sync'.
* 09:11 elukey@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'sync'.
* 09:10 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 09:10 elukey@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 09:09 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 09:09 elukey@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 08:55 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 08:54 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 08:40 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:40 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:36 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqsin [reason: depooling for maintainance, [[phab:T438052|T438052]]]
* 08:36 slyngshede@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool eqsin [reason: depooling for maintainance, [[phab:T438052|T438052]]]
* 08:35 slyngshede@cumin1003: END (FAIL) - Cookbook sre.dns.admin (exit_code=99) DNS admin: depool eqsin [reason: no reason specified, no task ID specified]
* 08:35 slyngshede@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool eqsin [reason: no reason specified, no task ID specified]
* 08:35 slyngshede@cumin1003: conftool action : set/pooled=no; selector: cluster=dnsbox,dc=eqsin
* 08:34 fabfur: start depooling eqsin ([[phab:T438052|T438052]])
* 08:24 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.20 refs [[phab:T430839|T430839]]
* 08:22 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:22 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:14 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.20 refs [[phab:T430839|T430839]]
* 08:11 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 08:11 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 08:11 jnuche@deploy1003: Rolling back deployment
* 08:10 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 08:07 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 07:59 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 07:59 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 07:58 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 07:54 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 07:34 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply
* 07:34 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply
* 07:20 mlitn@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342112{{!}}Adds an instrument for pre-image-carousel-retest (T437076)]], [[gerrit:1342113{{!}}Adds an instrument for pre-image-carousel-retest (T437076)]], [[gerrit:1342117{{!}}Set up instrument for 5-arm test (T437076)]], [[gerrit:1342118{{!}}Set up instrument for 5-arm test (T437076)]] (duration: 10m 56s)
* 07:16 mlitn@deploy1003: mlitn: Continuing with deployment
* 07:15 mlitn@deploy1003: mlitn: Backport for [[gerrit:1342112{{!}}Adds an instrument for pre-image-carousel-retest (T437076)]], [[gerrit:1342113{{!}}Adds an instrument for pre-image-carousel-retest (T437076)]], [[gerrit:1342117{{!}}Set up instrument for 5-arm test (T437076)]], [[gerrit:1342118{{!}}Set up instrument for 5-arm test (T437076)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be veri
* 07:09 mlitn@deploy1003: Started scap sync-world: Backport for [[gerrit:1342112{{!}}Adds an instrument for pre-image-carousel-retest (T437076)]], [[gerrit:1342113{{!}}Adds an instrument for pre-image-carousel-retest (T437076)]], [[gerrit:1342117{{!}}Set up instrument for 5-arm test (T437076)]], [[gerrit:1342118{{!}}Set up instrument for 5-arm test (T437076)]]
* 06:50 moritzm: installing sudo security updates
* 06:47 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply
* 06:37 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply
* 05:12 moritzm: pruned obsolete Bullseye image buildkitd from the docker registry [[phab:T416452|T416452]]
* 04:56 kart_: Updated Apertium to 2026-09-15-084320-production ([[phab:T437213|T437213]])
* 04:54 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/apertium: apply
* 04:54 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/apertium: apply
* 04:50 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/apertium: apply
* 04:49 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/apertium: apply
* 04:45 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/apertium: apply
* 04:45 kartik@deploy1003: helmfile [staging] START helmfile.d/services/apertium: apply
* 04:24 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 04:24 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 04:22 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 04:22 oblivian@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 36s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-15 ==
* 23:09 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply
* 23:08 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply
* 23:08 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply
* 23:08 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply
* 23:07 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 23:07 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 23:07 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 23:07 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 23:06 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 23:06 rzl@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 22:57 sukhe@puppetserver1001: conftool action : set/weight=1; selector: name=cp6001.drmrs.wmnet,service=cdn
* 22:50 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc-wf1002.eqiad.wmnet with OS trixie
* 22:46 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs ([[phab:T436363|T436363]])
* 22:46 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs6003.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:46 brett@cumin2003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs6003.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:46 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs6003.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:46 brett@cumin2003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs6003.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:45 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs6002.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:45 brett@cumin2003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs6002.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:45 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs6002.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:44 brett@cumin2003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs6002.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:44 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs6001.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:44 brett@cumin2003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs6001.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:43 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs6001.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:43 brett@cumin2003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs6001.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:43 brett@cumin2003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs ([[phab:T436363|T436363]])
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc-wf1002.eqiad.wmnet with reason: host reimage
* 22:33 brett@puppetserver1001: conftool action : set/weight=100; selector: name=cp6001.*
* 22:32 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp6001.*
* 22:26 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc-wf1002.eqiad.wmnet with reason: host reimage
* 22:07 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc-wf1002
* 22:07 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-wf1002
* 22:07 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host mc-wf1002
* 22:07 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc-wf1002.eqiad.wmnet 142.48.64.10.in-addr.arpa 2.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:07 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache mc-wf1002.eqiad.wmnet 142.48.64.10.in-addr.arpa 2.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:07 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:07 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc-wf1002 - rzl@cumin2003"
* 22:07 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc-wf1002 - rzl@cumin2003"
* 22:02 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:01 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host mc-wf1002
* 22:01 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host mc-wf1002.eqiad.wmnet with OS trixie
* 21:57 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6001.drmrs.wmnet with OS trixie
* 21:55 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply
* 21:55 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply
* 21:53 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply
* 21:53 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply
* 21:53 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:53 rzl@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:52 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:52 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:48 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:48 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:34 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6001.drmrs.wmnet with reason: host reimage
* 21:30 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp6001.drmrs.wmnet with reason: host reimage
* 21:20 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342049{{!}}MobileFrontend: Add app icons (T434258)]] (duration: 11m 47s)
* 21:15 jdlrobson@deploy1003: jdlrobson, cklimas: Continuing with deployment
* 21:13 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp6001.drmrs.wmnet with OS trixie
* 21:12 jdlrobson@deploy1003: jdlrobson, cklimas: Backport for [[gerrit:1342049{{!}}MobileFrontend: Add app icons (T434258)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:12 brett@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp6001.mgmt.drmrs.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:08 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1342049{{!}}MobileFrontend: Add app icons (T434258)]]
* 20:51 cdobbins@puppetserver1001: conftool action : set/weight=1; selector: name=cp2046.codfw.wmnet
* 20:51 cdobbins@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp2046.codfw.wmnet
* 20:48 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341271{{!}}Parsoid Read Views: Enable on all namespaces on wikitech (labswiki) (T437916)]] (duration: 09m 11s)
* 20:47 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp2046.codfw.wmnet with OS trixie
* 20:43 arlolra@deploy1003: ssastry, arlolra: Continuing with deployment
* 20:42 arlolra@deploy1003: ssastry, arlolra: Backport for [[gerrit:1341271{{!}}Parsoid Read Views: Enable on all namespaces on wikitech (labswiki) (T437916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:38 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1341271{{!}}Parsoid Read Views: Enable on all namespaces on wikitech (labswiki) (T437916)]]
* 20:34 brett@cumin2003: START - Cookbook sre.hosts.provision for host cp6001.mgmt.drmrs.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:30 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp6001.*
* 20:24 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp2046.codfw.wmnet with reason: host reimage
* 20:23 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339741{{!}}Enable ReaderExperiments in eswiki, jawiki, and ptwiki (T438009)]] (duration: 15m 58s)
* 20:20 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cp2046.codfw.wmnet with reason: host reimage
* 20:19 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T436363|T436363]])
* 20:19 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs4010.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:19 brett@cumin2003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs4010.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:19 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs4010.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:19 brett@cumin2003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs4010.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:18 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs4009.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:18 arlolra@deploy1003: lwatson, arlolra: Continuing with deployment
* 20:18 brett@cumin2003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs4009.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:17 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs4009.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:17 brett@cumin2003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs4009.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:17 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs4008.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:17 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp3075.esams.wmnet
* 20:17 sukhe@puppetserver1001: conftool action : set/weight=1; selector: name=cp3075.esams.wmnet
* 20:17 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp1103.eqiad.wmnet
* 20:17 brett@cumin2003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs4008.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:16 sukhe@puppetserver1001: conftool action : set/weight=1; selector: name=cp1103.eqiad.wmnet
* 20:16 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs4008.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:16 brett@cumin2003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs4008.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:16 brett@cumin2003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T436363|T436363]])
* 20:11 arlolra@deploy1003: lwatson, arlolra: Backport for [[gerrit:1339741{{!}}Enable ReaderExperiments in eswiki, jawiki, and ptwiki (T438009)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:10 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) config_reloading A:liberica-ulsfo ([[phab:T436363|T436363]])
* 20:08 brett@cumin2003: START - Cookbook sre.loadbalancer.admin config_reloading A:liberica-ulsfo ([[phab:T436363|T436363]])
* 20:08 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp1103.eqiad.wmnet with OS trixie
* 20:07 inflatador: bking@ganeti1046 sudo gnt-instance modify -B memory=4g,vcpus=4 dse-k8s-etcd100[1-3].eqiad.wmnet [[phab:T438084|T438084]]
* 20:07 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1339741{{!}}Enable ReaderExperiments in eswiki, jawiki, and ptwiki (T438009)]]
* 20:06 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp3075.esams.wmnet with OS trixie
* 20:04 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp7009.*
* 20:04 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host cp2046.codfw.wmnet with OS trixie
* 20:02 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp7009.*
* 20:02 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp7009.*
* 20:02 brett@puppetserver1001: conftool action : set/weight=1; selector: name=cp7009.*
* 20:01 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp7009.magru.wmnet with OS trixie
* 19:53 brett@puppetserver1001: conftool action : set/weight=1; selector: name=cp4045.*
* 19:53 brett@puppetserver1001: conftool action : set/weight=1; selector: name=cp4046.*
* 19:52 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4046.*
* 19:51 cdobbins@puppetserver1001: conftool action : set/pooled=no; selector: name=cp2046.codfw.wmnet
* 19:51 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp4046.ulsfo.wmnet with OS trixie
* 19:49 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4045.*
* 19:45 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp1103.eqiad.wmnet with reason: host reimage
* 19:43 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp4045.ulsfo.wmnet with OS trixie
* 19:41 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp3075.esams.wmnet with reason: host reimage
* 19:39 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cp1103.eqiad.wmnet with reason: host reimage
* 19:37 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp7009.magru.wmnet with reason: host reimage
* 19:33 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cp3075.esams.wmnet with reason: host reimage
* 19:32 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp7009.magru.wmnet with reason: host reimage
* 19:27 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp4046.ulsfo.wmnet with reason: host reimage
* 19:23 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp4046.ulsfo.wmnet with reason: host reimage
* 19:21 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp1103.eqiad.wmnet with OS trixie
* 19:19 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp1103.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:19 sukhe@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cp1103.eqiad.wmnet with reason: reimage
* 19:18 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp4045.ulsfo.wmnet with reason: host reimage
* 19:13 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 19:12 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 19:12 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 19:12 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp3075.esams.wmnet with OS trixie
* 19:12 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp4045.ulsfo.wmnet with reason: host reimage
* 19:12 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 19:11 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp3075.mgmt.esams.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:10 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp7009.magru.wmnet with OS trixie
* 19:09 brett@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp7009.mgmt.magru.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:08 sukhe@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cp3075.esams.wmnet with reason: reimaging
* 19:08 sukhe@cumin1004: START - Cookbook sre.hosts.provision for host cp1103.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:05 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp4046.ulsfo.wmnet with OS trixie
* 19:05 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore1006.eqiad.wmnet with OS bookworm
* 19:04 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3075.*
* 19:02 brett@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp4046.mgmt.ulsfo.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:01 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp1103.eqiad.wmnet
* 19:01 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp1103.eqiad.wmnet
* 19:00 sukhe@cumin1004: START - Cookbook sre.hosts.provision for host cp3075.mgmt.esams.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:58 brett@cumin2003: START - Cookbook sre.hosts.provision for host cp7009.mgmt.magru.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:57 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3075.esams.wmnet
* 18:55 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp1101.eqiad.wmnet
* 18:55 sukhe@puppetserver1001: conftool action : set/weight=1; selector: name=cp1101.eqiad.wmnet
* 18:55 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp4045.ulsfo.wmnet with OS trixie
* 18:52 brett@cumin2003: START - Cookbook sre.hosts.provision for host cp4046.mgmt.ulsfo.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:52 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp1101.eqiad.wmnet with OS trixie
* 18:45 cdobbins@puppetserver1001: conftool action : set/weight=1; selector: name=cp2044.codfw.wmnet
* 18:44 cdobbins@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp2044.codfw.wmnet
* 18:44 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore1006.eqiad.wmnet with reason: host reimage
* 18:40 brett@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp4045.mgmt.ulsfo.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:40 eevans@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore1006.eqiad.wmnet with reason: host reimage
* 18:39 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp3074.esams.wmnet
* 18:36 sukhe@cumin1004: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) config_reloading P<nowiki>{</nowiki>lvs3008.esams.wmnet<nowiki>}</nowiki> and A:liberica
* 18:36 sukhe@cumin1004: START - Cookbook sre.loadbalancer.admin config_reloading P<nowiki>{</nowiki>lvs3008.esams.wmnet<nowiki>}</nowiki> and A:liberica
* 18:33 sukhe@puppetserver1001: conftool action : set/weight=1; selector: name=cp3074.esams.wmnet
* 18:32 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp2044.codfw.wmnet with OS trixie
* 18:31 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp3074.esams.wmnet with OS trixie
* 18:30 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp1101.eqiad.wmnet with reason: host reimage
* 18:29 brett@cumin2003: START - Cookbook sre.hosts.provision for host cp4045.mgmt.ulsfo.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:26 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cp1101.eqiad.wmnet with reason: host reimage
* 18:22 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp7009.magru.wmnet with OS trixie
* 18:20 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1006.eqiad.wmnet with OS bookworm
* 18:20 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:19 eevans@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:19 eevans@cumin1004: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore1006.eqiad.wmnet
* 18:19 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore1006.eqiad.wmnet
* 18:10 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp1101.eqiad.wmnet with OS trixie
* 18:09 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp2044.codfw.wmnet with reason: host reimage
* 18:09 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp4045.ulsfo.wmnet with OS trixie
* 18:08 eevans@cumin1004: START - Cookbook sre.hosts.reboot-single for host sessionstore1006.eqiad.wmnet
* 18:07 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp3074.esams.wmnet with reason: host reimage
* 17:52 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp7009.magru.wmnet with reason: host reimage
* 17:48 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp2044.codfw.wmnet with OS trixie
* 17:43 cdobbins@puppetserver1001: conftool action : set/pooled=no; selector: name=cp2044.codfw.wmnet
* 17:36 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp3074.esams.wmnet with OS trixie
* 17:34 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp4046.ulsfo.wmnet with OS trixie
* 17:34 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp4045.ulsfo.wmnet with OS trixie
* 17:33 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3074.esams.wmnet
* 17:28 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp4046.*
* 17:28 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp4045.*
* 17:26 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp7009.magru.wmnet with OS trixie
* 17:25 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp1101.eqiad.wmnet with OS trixie
* 17:24 cdobbins@puppetserver1001: conftool action : set/pooled=no; selector: name=cp7009.magru.wmnet
* 17:23 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp1101.eqiad.wmnet
* 17:22 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:22 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org
* 17:02 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp7009.*
* 17:01 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:01 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:00 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:59 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:59 vriley@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host ms-be1099
* 16:59 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ms-be1099
* 16:59 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:56 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 16:55 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:55 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:55 vriley@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host ms-be1099
* 16:55 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ms-be1099
* 16:55 vriley@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host ms-be1099
* 16:54 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ms-be1099
* 16:54 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1098.eqiad.wmnet with OS bullseye
* 16:53 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:53 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [ms-be1099] - vriley@cumin1003"
* 16:53 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [ms-be1099] - vriley@cumin1003"
* 16:49 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 16:33 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS bullseye
* 16:21 mutante: temp disabling puppet on C:zookeeper (32 hosts) - safe deploy of https://gerrit.wikimedia.org/r/c/operations/puppet/+/1327569
* 16:04 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341890{{!}}Restore table borders for client-side MathJax (T435274)]], [[gerrit:1340558{{!}}lift IP cap for edit-a-thon /workshop (T437609 T437594 T437470)]] (duration: 24m 19s)
* 15:59 krinkle@deploy1003: anzx, krinkle: Continuing with deployment
* 15:44 krinkle@deploy1003: anzx, krinkle: Backport for [[gerrit:1341890{{!}}Restore table borders for client-side MathJax (T435274)]], [[gerrit:1340558{{!}}lift IP cap for edit-a-thon /workshop (T437609 T437594 T437470)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:40 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1341890{{!}}Restore table borders for client-side MathJax (T435274)]], [[gerrit:1340558{{!}}lift IP cap for edit-a-thon /workshop (T437609 T437594 T437470)]]
* 15:34 elukey@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 15:34 elukey@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 15:33 brennen@deploy1003: Finished deploy [phabricator/deployment@c386249]: deploy phab1005 for [[phab:T437930|T437930]] (duration: 00m 39s)
* 15:33 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS bullseye
* 15:33 brennen@deploy1003: Started deploy [phabricator/deployment@c386249]: deploy phab1005 for [[phab:T437930|T437930]]
* 15:32 elukey@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 15:32 elukey@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'sync'.
* 15:32 brennen@deploy1003: Finished deploy [phabricator/deployment@c386249]: deploy phab2003 for [[phab:T437930|T437930]] (duration: 00m 52s)
* 15:32 elukey@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 15:32 elukey@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'sync'.
* 15:31 brennen@deploy1003: Started deploy [phabricator/deployment@c386249]: deploy phab2003 for [[phab:T437930|T437930]]
* 15:31 elukey@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'.
* 15:31 elukey@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'.
* 15:26 jelto@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on phab2003.codfw.wmnet,phab[1005-1006].eqiad.wmnet with reason: Phabricator deploy
* 15:26 moritzm: pruned obsolete Bullseye image amd-gpu-tester from the docker registry [[phab:T416452|T416452]]
* 15:12 elukey@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 15:12 elukey@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 15:11 elukey@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 15:11 elukey@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 15:00 tgr@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341292{{!}}CommonSettings: Use a restrictive CSP for auth.wikimedia.org (T419684)]] (duration: 25m 11s)
* 14:55 tgr@deploy1003: tgr, arendpieter: Continuing with deployment
* 14:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:53 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:53 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:52 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:49 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 14:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 14:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 14:47 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 14:47 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 14:45 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 14:45 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 14:45 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 14:44 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 14:42 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 14:42 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 14:39 tgr@deploy1003: tgr, arendpieter: Backport for [[gerrit:1341292{{!}}CommonSettings: Use a restrictive CSP for auth.wikimedia.org (T419684)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:34 tgr@deploy1003: Started scap sync-world: Backport for [[gerrit:1341292{{!}}CommonSettings: Use a restrictive CSP for auth.wikimedia.org (T419684)]]
* 14:17 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341895{{!}}ReportIncidentController: Instance cache expensive methods (T437588)]] (duration: 11m 56s)
* 14:16 btullis@cumin1004: END (PASS) - Cookbook sre.ceph.rotate-osd-keys (exit_code=0) rolling rotate_keys on A:cephosd-codfw
* 14:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:12 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 14:09 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1341895{{!}}ReportIncidentController: Instance cache expensive methods (T437588)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:05 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1341895{{!}}ReportIncidentController: Instance cache expensive methods (T437588)]]
* 13:43 btullis@cumin1004: START - Cookbook sre.ceph.rotate-osd-keys rolling rotate_keys on A:cephosd-codfw
* 13:36 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341861{{!}}SuggestedInvestigations: Update "sockpuppet" queue view defaults (T438018)]] (duration: 10m 23s)
* 13:32 stran@deploy1003: stran: Continuing with deployment
* 13:30 stran@deploy1003: stran: Backport for [[gerrit:1341861{{!}}SuggestedInvestigations: Update "sockpuppet" queue view defaults (T438018)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:26 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1341861{{!}}SuggestedInvestigations: Update "sockpuppet" queue view defaults (T438018)]]
* 13:21 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:20 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:19 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334946{{!}}ArticleGuidance: Remove the experiment configuration keys (T434487)]] (duration: 09m 19s)
* 13:16 btullis@cumin1004: END (PASS) - Cookbook sre.ceph.rotate-osd-keys (exit_code=0) rolling rotate_keys on P<nowiki>{</nowiki>cephosd2001.codfw.wmnet<nowiki>}</nowiki> and (A:cephosd-codfw or A:cephosd-eqiad)
* 13:15 sbisson@deploy1003: sbisson: Continuing with deployment
* 13:14 sbisson@deploy1003: sbisson: Backport for [[gerrit:1334946{{!}}ArticleGuidance: Remove the experiment configuration keys (T434487)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) config_reloading P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:10 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1334946{{!}}ArticleGuidance: Remove the experiment configuration keys (T434487)]]
* 13:10 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin config_reloading P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:08 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:08 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:07 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:07 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:07 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:07 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:03 btullis@cumin1004: START - Cookbook sre.ceph.rotate-osd-keys rolling rotate_keys on P<nowiki>{</nowiki>cephosd2001.codfw.wmnet<nowiki>}</nowiki> and (A:cephosd-codfw or A:cephosd-eqiad)
* 13:00 btullis@cumin1004: END (PASS) - Cookbook sre.ceph.rotate-osd-keys (exit_code=0) rolling rotate_keys on P<nowiki>{</nowiki>cephosd2001.codfw.wmnet<nowiki>}</nowiki> and (A:cephosd-codfw or A:cephosd-eqiad)
* 12:59 btullis@cumin1004: START - Cookbook sre.ceph.rotate-osd-keys rolling rotate_keys on P<nowiki>{</nowiki>cephosd2001.codfw.wmnet<nowiki>}</nowiki> and (A:cephosd-codfw or A:cephosd-eqiad)
* 12:46 btullis@cumin1004: END (PASS) - Cookbook sre.ceph.rotate-osd-keys (exit_code=0) rolling rotate_keys on P<nowiki>{</nowiki>cephosd2001.codfw.wmnet<nowiki>}</nowiki> and (A:cephosd-codfw or A:cephosd-eqiad)
* 12:45 btullis@cumin1004: START - Cookbook sre.ceph.rotate-osd-keys rolling rotate_keys on P<nowiki>{</nowiki>cephosd2001.codfw.wmnet<nowiki>}</nowiki> and (A:cephosd-codfw or A:cephosd-eqiad)
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: network maintenance finished, [[phab:T437984|T437984]]]
* 12:40 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: network maintenance finished, [[phab:T437984|T437984]]]
* 12:29 XioNoX: asw1-b4-magru> request system reboot - [[phab:T437984|T437984]]
* 12:24 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 12:24 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 12:24 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 12:23 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 12:23 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 12:23 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 12:22 moritzm: installing shadow security updates
* 12:19 slyngshede@puppetserver1001: conftool action : set/weight=1; selector: name=cp7010.magru.wmnet
* 12:13 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 12:13 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 12 hosts with reason: Switch maintenance
* 12:12 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-b4-magru,asw1-b4-magru IPv6,asw1-b4-magru.mgmt with reason: Switch maintenance
* 12:11 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 12:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: switch reboot, [[phab:T437984|T437984]]]
* 12:11 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: switch reboot, [[phab:T437984|T437984]]]
* 12:09 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on install7002.wikimedia.org with reason: switch reboot
* 12:08 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 12:07 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 12:07 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 12:07 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 12:06 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 12:06 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 12:03 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 12:03 XioNoX: push pfw policies - [[phab:T437627|T437627]]
* 12:01 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 12:01 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) config_reloading P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 12:00 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin config_reloading P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 11:56 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 11:56 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 11:33 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2250.codfw.wmnet with reason: cloning db2201
* 11:18 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7004.magru.wmnet
* 11:17 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7004.magru.wmnet
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 11:05 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 11:05 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 11:04 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 10:51 slyngshede@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp7010.magru.wmnet
* 10:34 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 10:33 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 10:30 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1098.eqiad.wmnet with OS trixie
* 10:21 moritzm: failover Ganeti master in magru to ganeti7001
* 10:20 moritzm: increased DRBD replication speed in Ganeti/magru [[phab:T428878|T428878]]
* 10:10 hashar@deploy1003: Finished deploy [integration/docroot@5cf09c8]: build: Updating npm dependencies (duration: 00m 13s)
* 10:10 hashar@deploy1003: Started deploy [integration/docroot@5cf09c8]: build: Updating npm dependencies
* 10:09 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 10:08 moritzm: increased DRBD replication speed in Ganeti/esams [[phab:T428878|T428878]]
* 10:07 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 10:05 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 10:05 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:39 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: switches reboot, [[phab:T437984|T437984]]]
* 09:39 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool esams [reason: switches reboot, [[phab:T437984|T437984]]]
* 09:31 XioNoX: asw1-by27-esams> request system reboot - [[phab:T437984|T437984]]
* 09:30 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS trixie
* 09:28 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp7010.magru.wmnet with OS trixie
* 09:26 ayounsi@cumin1003: END (FAIL) - Cookbook sre.network.depool-rack (exit_code=99) with action 'depool' for esams rack BY27
* 09:24 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for esams rack BY27
* 09:24 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1098.eqiad.wmnet with OS trixie
* 09:23 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS trixie
* 09:22 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be1098.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 09:15 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.20 refs [[phab:T430839|T430839]]
* 09:10 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1098.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 09:06 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-by27-esams,asw1-by27-esams IPv6,asw1-by27-esams.mgmt with reason: Switch maintenance
* 09:05 ayounsi@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on asw1-by27-esams IPv6,asw1-by27-esams.mgmt,asw1-by-27-esams with reason: Switch maintenance
* 09:04 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 12 hosts with reason: Switch maintenance
* 09:04 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp7010.magru.wmnet with reason: host reimage
* 09:01 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: switches reboot, [[phab:T437984|T437984]]]
* 09:00 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: switches reboot, [[phab:T437984|T437984]]]
* 09:00 slyngshede@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp7010.magru.wmnet with reason: host reimage
* 08:59 jnuche@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341698{{!}}RestSandbox: Pass JsonLocalizer instead of ResponseFactory to ModuleManager (T437982)]] (duration: 12m 03s)
* 08:55 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2197.codfw.wmnet with reason: cloning db2201
* 08:55 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on install3004.wikimedia.org with reason: switch reboot
* 08:53 jnuche@deploy1003: jnuche: Continuing with deployment
* 08:52 jnuche@deploy1003: jnuche: Backport for [[gerrit:1341698{{!}}RestSandbox: Pass JsonLocalizer instead of ResponseFactory to ModuleManager (T437982)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:49 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: sync
* 08:49 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: sync
* 08:47 jnuche@deploy1003: Started scap sync-world: Backport for [[gerrit:1341698{{!}}RestSandbox: Pass JsonLocalizer instead of ResponseFactory to ModuleManager (T437982)]]
* 08:33 slyngshede@cumin1003: START - Cookbook sre.hosts.reimage for host cp7010.magru.wmnet with OS trixie
* 08:26 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be1098.eqiad.wmnet with OS trixie
* 08:26 slyngshede@puppetserver1001: conftool action : set/pooled=no; selector: name=cp7010.magru.wmnet
* 08:25 XioNoX: asw1-b3-magru - Disable logging and file logging for BRCM_PKT - [[phab:T437984|T437984]]
* 08:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS trixie
* 08:18 dpogorzelski@dns1004: END - running authdns-update
* 08:15 dpogorzelski@dns1004: START - running authdns-update
* 08:14 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3005.esams.wmnet
* 08:13 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3005.esams.wmnet
* 08:07 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341274{{!}}SI: Implement "queue view" functionality (T437183)]], [[gerrit:1341242{{!}}SuggestedInvestigations: Add and enable 'sockpuppets' queue view (T437183)]], [[gerrit:1341254{{!}}Add wmf-specific Special:SuggestedInvestigations messages (T437183)]] (duration: 55m 27s)
* 07:54 stran@deploy1003: stran: Continuing with deployment
* 07:31 stran@deploy1003: stran: Backport for [[gerrit:1341274{{!}}SI: Implement "queue view" functionality (T437183)]], [[gerrit:1341242{{!}}SuggestedInvestigations: Add and enable 'sockpuppets' queue view (T437183)]], [[gerrit:1341254{{!}}Add wmf-specific Special:SuggestedInvestigations messages (T437183)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:18 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply
* 07:18 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply
* 07:11 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1341274{{!}}SI: Implement "queue view" functionality (T437183)]], [[gerrit:1341242{{!}}SuggestedInvestigations: Add and enable 'sockpuppets' queue view (T437183)]], [[gerrit:1341254{{!}}Add wmf-specific Special:SuggestedInvestigations messages (T437183)]]
* 07:06 moritzm: pruned obsolete Bullseye image python3-devel from the docker registry [[phab:T416452|T416452]]
* 06:51 moritzm: pruned obsolete Bullseye image python3-build-bullseye from the docker registry [[phab:T416452|T416452]]
* 05:59 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply
* 05:49 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply
* 05:48 oblivian@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 05:47 oblivian@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 05:38 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply
* 05:28 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-video: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-video: apply
* 05:10 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-video: apply
* 05:10 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-video: apply
* 05:10 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-video: apply
* 05:10 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-video: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-syntaxhighlight: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-syntaxhighlight: apply
* 05:10 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-syntaxhighlight: apply
* 05:10 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-syntaxhighlight: apply
* 05:10 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-syntaxhighlight: apply
* 05:10 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-syntaxhighlight: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-media: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-media: apply
* 05:10 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-media: apply
* 05:10 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-media: apply
* 05:10 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-media: apply
* 05:10 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-media: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-constraints: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-constraints: apply
* 05:10 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-constraints: apply
* 05:09 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-constraints: apply
* 05:09 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-constraints: apply
* 05:09 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-constraints: apply
* 05:08 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply
* 05:08 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox: apply
* 05:07 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply
* 05:07 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox: apply
* 05:07 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox: apply
* 05:07 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox: apply
* 05:07 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply
* 05:07 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply
* 05:06 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-timeline: apply
* 05:06 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-timeline: apply
* 05:06 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply
* 05:06 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply
* 04:07 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.17 (duration: 07m 10s)
* 03:06 mwpresync@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.19,1.47.0-wmf.20,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.20 refs [[phab:T430839|T430839]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 22s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:43 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 00:23 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore1005.eqiad.wmnet with reason: host reimage
* 00:19 jclark@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore1005.eqiad.wmnet with reason: host reimage
* 00:17 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 00:07 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sessionstore1005.eqiad.wmnet with OS bookworm
== 2026-09-14 ==
* 23:41 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 23:26 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324966{{!}}Enable Produnto on pilot wikis (T421436)]] (duration: 12m 59s)
* 23:22 tstarling@deploy1003: tstarling: Continuing with deployment
* 23:17 tstarling@deploy1003: tstarling: Backport for [[gerrit:1324966{{!}}Enable Produnto on pilot wikis (T421436)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:13 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1324966{{!}}Enable Produnto on pilot wikis (T421436)]]
* 23:01 eevans@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 22:41 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341385{{!}}VisualEditor: don't register settings tool in wikitextCommandRegistry (T437810)]] (duration: 09m 22s)
* 22:36 kemayo@deploy1003: kemayo: Continuing with deployment
* 22:36 kemayo@deploy1003: kemayo: Backport for [[gerrit:1341385{{!}}VisualEditor: don't register settings tool in wikitextCommandRegistry (T437810)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:31 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1341385{{!}}VisualEditor: don't register settings tool in wikitextCommandRegistry (T437810)]]
* 22:22 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 22:22 eevans@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 21:43 sbassett: Deployed security fix for [[phab:T435623|T435623]]
* 21:29 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1098.eqiad.wmnet with OS bullseye
* 21:29 sbassett: Deployed security fix for [[phab:T434372|T434372]]
* 21:26 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 21:26 eevans@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 21:05 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338999{{!}}Enable ReadingLists for all logged-in users on English Wikipedia (T434923)]], [[gerrit:1340004{{!}}Enable Reading Recommendations experiment on test wiki (T437665)]] (duration: 11m 03s)
* 21:00 aude@deploy1003: aude, jdlrobson: Continuing with deployment
* 20:58 aude@deploy1003: aude, jdlrobson: Backport for [[gerrit:1338999{{!}}Enable ReadingLists for all logged-in users on English Wikipedia (T434923)]], [[gerrit:1340004{{!}}Enable Reading Recommendations experiment on test wiki (T437665)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:54 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1338999{{!}}Enable ReadingLists for all logged-in users on English Wikipedia (T434923)]], [[gerrit:1340004{{!}}Enable Reading Recommendations experiment on test wiki (T437665)]]
* 20:47 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341278{{!}}[A11y] Add list semantics to ReadingList page (T435864 T434923)]] (duration: 12m 49s)
* 20:43 aude@deploy1003: aude, jdlrobson: Continuing with deployment
* 20:39 aude@deploy1003: aude, jdlrobson: Backport for [[gerrit:1341278{{!}}[A11y] Add list semantics to ReadingList page (T435864 T434923)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:34 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1341278{{!}}[A11y] Add list semantics to ReadingList page (T435864 T434923)]]
* 20:32 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339811{{!}}wmf-config: Register content/v2-beta REST module as disabled (T432798)]], [[gerrit:1338274{{!}}wikifunctions: Move abstract fragments to mainstash (T432849)]] (duration: 25m 21s)
* 20:28 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:27 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:27 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:27 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:25 jforrester@deploy1003: jforrester, aghirelli: Continuing with deployment
* 20:24 jforrester@deploy1003: jforrester, aghirelli: Backport for [[gerrit:1339811{{!}}wmf-config: Register content/v2-beta REST module as disabled (T432798)]], [[gerrit:1338274{{!}}wikifunctions: Move abstract fragments to mainstash (T432849)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS bullseye
* 20:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be1098.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 20:06 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1339811{{!}}wmf-config: Register content/v2-beta REST module as disabled (T432798)]], [[gerrit:1338274{{!}}wikifunctions: Move abstract fragments to mainstash (T432849)]]
* 20:04 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1098.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 20:03 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be1098
* 20:02 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ms-be1098
* 20:02 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 20:02 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [ms-be1098] - vriley@cumin1003"
* 20:02 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [ms-be1098] - vriley@cumin1003"
* 19:59 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 19:58 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 19:57 dzahn@dns1005: END - running authdns-update
* 19:55 eevans@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 19:55 dzahn@dns1005: START - running authdns-update
* 19:54 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 19:38 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341273{{!}}[CodeMirror] enable for new users (enwiki), new VE integration (global) (T288161 T432558)]] (duration: 33m 51s)
* 19:26 musikanimal@deploy1003: musikanimal: Continuing with deployment
* 19:22 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1341273{{!}}[CodeMirror] enable for new users (enwiki), new VE integration (global) (T288161 T432558)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:04 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1341273{{!}}[CodeMirror] enable for new users (enwiki), new VE integration (global) (T288161 T432558)]]
* 18:53 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:50 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 18:50 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 18:49 eevans@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 18:26 brett@cumin2003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d6-eqiad
* 18:26 brett@cumin2003: START - Cookbook sre.network.tls for network device lsw1-d6-eqiad
* 18:26 brett@cumin2003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-eqiad
* 18:26 brett@cumin2003: START - Cookbook sre.network.tls for network device ssw1-d8-eqiad
* 18:25 brett@cumin2003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-eqiad
* 18:25 brett@cumin2003: START - Cookbook sre.network.tls for network device lsw1-d4-eqiad
* 18:25 root@cumin2003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-eqiad
* 18:25 root@cumin2003: START - Cookbook sre.network.tls for network device lsw1-d2-eqiad
* 18:17 jclark@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:14 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337611{{!}}extension-list: Add ModeratorToolkit (T431000)]] (duration: 09m 34s)
* 18:10 samtar@deploy1003: samtar: Continuing with deployment
* 18:09 samtar@deploy1003: samtar: Backport for [[gerrit:1337611{{!}}extension-list: Add ModeratorToolkit (T431000)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:06 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 18:05 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1337611{{!}}extension-list: Add ModeratorToolkit (T431000)]]
* 18:01 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 17:48 jclark@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 17:07 tgr@deploy1003: Finished scap sync-world: Backport for [[gerrit:1330446{{!}}CommonSettings: Use a restrictive, eval-free CSP for auth.wikimedia.org (T419684)]] (duration: 23m 19s)
* 17:00 tgr@deploy1003: arendpieter, tgr: Rolling back deployment
* 16:52 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 16:51 eevans@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 16:49 tgr@deploy1003: arendpieter, tgr: Backport for [[gerrit:1330446{{!}}CommonSettings: Use a restrictive, eval-free CSP for auth.wikimedia.org (T419684)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:44 tgr@deploy1003: Started scap sync-world: Backport for [[gerrit:1330446{{!}}CommonSettings: Use a restrictive, eval-free CSP for auth.wikimedia.org (T419684)]]
* 16:09 Amir1: drop links tables from db1252 ([[phab:T437278|T437278]])
* 16:07 Amir1: drop links tables from db2240 ([[phab:T437278|T437278]])
* 16:05 Amir1: drop non-links tables from db2247 ([[phab:T437278|T437278]])
* 15:53 Lucas_WMDE: UTC afternoon backport+config window belatedly done
* 15:50 lucaswerkmeister-wmde@deploy1003: mwscript-k8s job started: namespaceDupes abstractwiki --fix # [[phab:T437772|T437772]]
* 15:49 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335401{{!}}Adjust extendedconfirmed calculation to first edit on viwiki (T437006)]], [[gerrit:1340505{{!}}core-Namespaces: Add AW and AWT alias for its talk in abstractwiki (T437772)]] (duration: 10m 23s)
* 15:48 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 15:47 eevans@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 15:45 lucaswerkmeister-wmde@deploy1003: bunnypranav, lucaswerkmeister-wmde, tryvix1509: Continuing with deployment
* 15:43 lucaswerkmeister-wmde@deploy1003: bunnypranav, lucaswerkmeister-wmde, tryvix1509: Backport for [[gerrit:1335401{{!}}Adjust extendedconfirmed calculation to first edit on viwiki (T437006)]], [[gerrit:1340505{{!}}core-Namespaces: Add AW and AWT alias for its talk in abstractwiki (T437772)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:39 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1335401{{!}}Adjust extendedconfirmed calculation to first edit on viwiki (T437006)]], [[gerrit:1340505{{!}}core-Namespaces: Add AW and AWT alias for its talk in abstractwiki (T437772)]]
* 15:36 elukey@dns1004: END - running authdns-update
* 15:36 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 15:35 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 15:35 moritzm: installing shadow security updates
* 15:35 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 15:34 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:33 elukey@dns1004: START - running authdns-update
* 15:33 eevans@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:32 eevans@cumin1004: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore1005.eqiad.wmnet
* 15:29 lucaswerkmeister-wmde@deploy1003: mwscript-k8s job started: namespaceDupes afwiki --fix # [[phab:T437576|T437576]]
* 15:29 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338902{{!}}afwiki: Create Draft and Draft talk namespaces (T437576)]] (duration: 15m 44s)
* 15:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 15:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 15:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 15:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 15:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 15:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 15:21 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, tryvix1509: Continuing with deployment
* 15:21 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:17 eevans@cumin1004: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore1005.eqiad.wmnet
* 15:17 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, tryvix1509: Backport for [[gerrit:1338902{{!}}afwiki: Create Draft and Draft talk namespaces (T437576)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:17 eevans@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore1005.eqiad.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 15:13 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore1004.eqiad.wmnet
* 15:13 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1338902{{!}}afwiki: Create Draft and Draft talk namespaces (T437576)]]
* 15:06 eevans@cumin1004: START - Cookbook sre.hosts.reboot-single for host sessionstore1004.eqiad.wmnet
* 14:59 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore1004.eqiad.wmnet with OS bookworm
* 14:38 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore1004.eqiad.wmnet with reason: host reimage
* 14:33 marostegui@dns1004: END - running authdns-update
* 14:32 eevans@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore1004.eqiad.wmnet with reason: host reimage
* 14:30 marostegui@dns1004: START - running authdns-update
* 14:15 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1004.eqiad.wmnet with OS bookworm
* 14:14 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore1004.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:13 eevans@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1004.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:13 eevans@cumin1004: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore1004.eqiad.wmnet
* 14:13 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore1004.eqiad.wmnet
* 14:05 moritzm: kick off a rebuild of base images on build2004
* 14:05 moritzm: kick off a rebuild of base images on build2004
* 14:00 eevans@cumin1004: START - Cookbook sre.hosts.reboot-single for host sessionstore1004.eqiad.wmnet
* 14:00 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:00 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 13:56 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:54 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:52 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:44 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 13:43 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 13:42 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:40 eevans@cumin1004: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore1004.eqiad.wmnet
* 13:40 eevans@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore1004.eqiad.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 13:40 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:40 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:39 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:39 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:39 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 13:38 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 13:38 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:38 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:38 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 13:37 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 13:37 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:36 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:36 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 13:35 oblivian@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:35 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 13:35 oblivian@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:35 oblivian@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:34 oblivian@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'.
* 13:34 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 13:34 oblivian@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:34 oblivian@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:34 oblivian@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:34 oblivian@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 13:34 oblivian@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'.
* 13:33 oblivian@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'.
* 13:33 oblivian@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
* 13:33 oblivian@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
* 13:33 oblivian@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 13:33 oblivian@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 13:33 oblivian@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 13:32 oblivian@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 13:32 oblivian@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 13:32 oblivian@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 13:32 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 13:32 oblivian@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 13:32 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 13:32 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 13:29 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:23 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1003.eqiad.wmnet
* 13:21 sukhe: sudo cumin -b11 "A:cp-text" "run-puppet-agent --enable 'merging CR 1338134'" [[phab:T425441|T425441]]
* 13:20 sukhe: sudo cumin -b11 "A:cp-text" "run-puppet-agent --enable 'merging CR 1338134'"[[phab:T425441|T425441]]
* 13:19 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:17 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1003.eqiad.wmnet
* 13:14 moritzm: installing Bird security updates
* 13:09 sukhe: sudo cumin "A:cp-text" "disable-puppet 'merging CR 1338134'"
* 13:06 jmm@dns1004: END - running authdns-update
* 13:04 jmm@dns1004: START - running authdns-update
* 12:58 moritzm: update Trixie installer image to 13.7 [[phab:T437715|T437715]]
* 12:58 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 12:54 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 12:52 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 12:49 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 12:48 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 12:47 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 12:44 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 12:44 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 12:42 oblivian@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 12:40 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 12:40 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 12:40 oblivian@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 12:39 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 12:39 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 12:39 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 12:37 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 12:37 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 12:36 marostegui@cumin1004: conftool action : set/pooled=yes; selector: name=clouddb1025.eqiad.wmnet,service=x4
* 12:34 _joe_: adding gvisor labels to all wikikube clusters nodes
* 12:30 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 12:14 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 12:14 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 11:22 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1260: After cloning
* 10:48 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:45 ladsgroup@dns1004: END - running authdns-update
* 10:42 ladsgroup@dns1004: START - running authdns-update
* 10:37 marostegui@cumin1004: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=x4
* 10:37 marostegui@cumin1004: START - Cookbook sre.mysql.pool pool db1260: After cloning
* 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 10:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 10:04 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 10:04 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 09:53 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:53 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:41 Amir1: drop links tables from db2172 ([[phab:T437278|T437278]])
* 09:40 Amir1: drop links tables from db1228 ([[phab:T437278|T437278]])
* 09:08 marostegui: Stop mariadb on db1260 to clone dbstore1007, there will be lag on wikireplicas:x4 https://phabricator.wikimedia.org/T437839
* 09:07 taavi@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1025.eqiad.wmnet
* 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1260: Needs to clone another host from this one
* 09:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1260: Needs to clone another host from this one
* 09:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb[1024-1025].eqiad.wmnet,db[1155,1260].eqiad.wmnet,dbstore1007.eqiad.wmnet with reason: Adding x4
* 08:44 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:42 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:41 moritzm: pruned obsolete Bullseye images php8.3-icu72-cli / php8.3-icu72-fpm-multiversion-base / php8.3-icu72-fpm from the docker registry [[phab:T416452|T416452]]
* 08:37 taavi@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet
* 08:37 taavi@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1024.eqiad.wmnet
* 08:36 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on dbstore1007.eqiad.wmnet with reason: Adding x4
* 08:35 moritzm: pruned obsolete Bullseye images php8.1-cli/php8.1-fpm/ php8.1-fpm-multiversion-base from the docker registry [[phab:T416452|T416452]]
* 08:10 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: sync
* 08:08 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: sync
* 07:58 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 23s)
* 07:57 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 07:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 07:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 07:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 07:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 07:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 07:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 07:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 07:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 07:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 07:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 07:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 07:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 07:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 07:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 07:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 07:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 07:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 07:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 07:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 07:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 07:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 07:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 07:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 07:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 07:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 07:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 07:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 07:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 07:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 07:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 07:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 07:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 07:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 07:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 07:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 07:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 07:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 07:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 07:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 07:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 07:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 07:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 07:31 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1340801{{!}}Allow title-like strings with Package: prefix in require() (T430644)]], [[gerrit:1340802{{!}}Runtime: Add a facility for loading files by title (T430644)]] (duration: 34m 30s)
* 07:18 tstarling@deploy1003: tstarling: Continuing with deployment
* 07:17 tstarling@deploy1003: tstarling: Backport for [[gerrit:1340801{{!}}Allow title-like strings with Package: prefix in require() (T430644)]], [[gerrit:1340802{{!}}Runtime: Add a facility for loading files by title (T430644)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 06:56 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1340801{{!}}Allow title-like strings with Package: prefix in require() (T430644)]], [[gerrit:1340802{{!}}Runtime: Add a facility for loading files by title (T430644)]]
* 06:26 TimStarling: on deploy1003: docker image pull docker-registry.wikimedia.org/php8.3-fpm-multiversion-base
* 05:51 _joe_: pulled bookworm:latest from build2004 to build2001 [[phab:T437829|T437829]]
* 05:39 _joe_: force-running build-base-images on build2004 for [[phab:T437829|T437829]]
* 04:53 TimStarling: on build2001 rebuilding base images [[phab:T437829|T437829]]
* 03:00 tstarling@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.18,1.47.0-wmf.19,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted
* 02:59 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1340801{{!}}Allow title-like strings with Package: prefix in require() (T430644)]], [[gerrit:1340802{{!}}Runtime: Add a facility for loading files by title (T430644)]]
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-13 ==
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 29s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-12 ==
* 19:40 ladsgroup@cumin1003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-eqiad
* 19:32 ladsgroup@cumin1003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-eqiad
* 19:30 ladsgroup@cumin1003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw
* 19:21 ladsgroup@cumin1003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 35s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-11 ==
* 21:51 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 21:50 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 16:47 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339808{{!}}Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813{{!}}Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]] (duration: 07m 23s)
* 16:43 sbassett@deploy1003: sbassett: Continuing with deployment
* 16:42 sbassett@deploy1003: sbassett: Backport for [[gerrit:1339808{{!}}Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813{{!}}Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:40 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1339808{{!}}Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813{{!}}Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]]
* 16:08 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 14:39 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 14:10 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:10 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:10 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:09 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 13:40 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:28 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:11 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:10 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: db1199 repool
* 11:05 moritzm: installing Linux 6.1.187 on Bookworm hosts
* 11:05 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Jcrespo out of all services on: 2443 hosts
* 10:44 aokoth@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2020 in turn
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2019 in turn
* 10:41 Emperor: restart versitygw@objectstorage0[0-3].service on backup2018 in turn
* 10:40 Emperor: restart versitygw@objectstorage0[0-3].service on backup2017 in turn
* 10:39 Emperor: restart versitygw@objectstorage0[0-3].service on backup2016 in turn
* 10:37 Emperor: restart versitygw@objectstorage0[0-3].service on backup2015 in turn
* 10:36 Emperor: restart versitygw@objectstorage0[0-3].service on backup1020 in turn
* 10:35 Emperor: restart versitygw@objectstorage0[0-3].service on backup1019 in turn
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1199: db1199 repool
* 10:33 Emperor: restart versitygw@objectstorage0[0-3].service on backup1018 in turn
* 10:32 Emperor: restart versitygw@objectstorage0[0-3].service on backup1017 in turn
* 10:30 Emperor: restart versitygw@objectstorage0[0-3].service on backup1016 in turn
* 10:20 Emperor: restart versitygw@objectstorage0[1-3].service on backup1015 in turn
* 10:17 Emperor: restart versitygw@objectstorage00.service on backup1015
* 10:15 aokoth@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 08:46 slyngs: Update CAS/SSO to CAS 7.3.8.3
* 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:45 slyngshede@dns1004: END - running authdns-update
* 08:43 slyngshede@dns1004: START - running authdns-update
* 08:36 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:28 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 7 hosts with reason: Restarting s5
* 08:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db[1154,1269].eqiad.wmnet with reason: Restarting s5
* 08:20 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:20 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Repooling db1159
* 07:58 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:58 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:54 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:28 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:27 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1199: Needs to clone another host from this one
* 07:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1199: Needs to clone another host from this one
* 07:16 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Repooling db1159
* 07:15 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1199.eqiad.wmnet with reason: Cloning s4
* 07:10 TimStarling: killed jobs for [[phab:T437056|T437056]] since they weren't purging
* 06:38 TimStarling: also started refreshLinks for ptwiki and zhwiki, reparsing ~3000 pages altogether [[phab:T437056|T437056]]
* 06:27 TimStarling: for [[phab:T437056|T437056]]: mwscript-k8s refreshLinks.php --wiki=eswiki --tracking-category scribunto-common-error-category
* 05:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1159: Needs to clone another host from this one
* 05:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1159: Needs to clone another host from this one
* 05:30 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1159.eqiad.wmnet with reason: Cloning
* 05:29 marostegui: Start cloning db1245:s5 [[phab:T437563|T437563]]
* 05:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet,db1245.eqiad.wmnet with reason: Cloning
* 05:25 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] (duration: 09m 59s)
* 05:21 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:20 tstarling@deploy1003: tstarling: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:15 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 50s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-10 ==
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1349.eqiad.wmnet
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1349.eqiad.wmnet
* 23:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1349.eqiad.wmnet
* 23:09 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1349
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1349
* 22:32 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1349
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:28 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1349
* 22:27 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:27 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1349.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1348.eqiad.wmnet
* 22:23 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 22:12 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] (duration: 10m 59s)
* 22:06 musikanimal@deploy1003: kemayo, musikanimal: Rolling back deployment
* 22:05 musikanimal@deploy1003: kemayo, musikanimal: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:01 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]]
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:00 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:52 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:47 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:47 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] (duration: 13m 23s)
* 21:46 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:46 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2006.codfw.wmnet
* 21:42 derenrich@deploy1003: derenrich: Continuing with deployment
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1348
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1348
* 21:39 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1348
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:39 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:37 derenrich@deploy1003: derenrich: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:35 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:34 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1348
* 21:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 21:33 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1348.eqiad.wmnet
* 21:33 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]]
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1348.eqiad.wmnet
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1348.eqiad.wmnet
* 21:31 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2006.codfw.wmnet
* 21:31 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] (duration: 09m 45s)
* 21:27 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:26 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]]
* 21:19 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:19 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2006.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 21:17 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] (duration: 13m 54s)
* 21:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2005.codfw.wmnet with OS bookworm
* 21:12 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:07 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdeb
* 21:03 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]]
* 20:54 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:52 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] (duration: 23m 49s)
* 20:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:47 jdrewniak@deploy1003: jdrewniak, milazg: Continuing with deployment
* 20:34 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1346.eqiad.wmnet
* 20:32 jdrewniak@deploy1003: jdrewniak, milazg: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2005.codfw.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:28 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]]
* 20:27 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2005.codfw.wmnet
* 20:26 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:24 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:22 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] (duration: 11m 24s)
* 20:17 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 20:15 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2005.codfw.wmnet
* 20:14 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:12 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 20:11 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:07 jdrewniak@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.18,1.47.0-wmf.19,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted
* 20:05 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:02 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:01 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2005.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 19:53 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:50 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1346
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 19:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1346.eqiad.wmnet
* 19:21 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 thcipriani: Gerrit downtime incoming for upgrade
* 19:17 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hadoop.roll-restart-workers (exit_code=0) restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 dzahn@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:30:00 on gerrit.wikimedia.org with reason: maintenance upgrade
* 19:16 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on gerrit2003.wikimedia.org with reason: maintenance upgrade
* 19:04 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 18:21 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e] (duration: 00m 59s)
* 18:20 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e]
* 18:19 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e] (duration: 05m 13s)
* 18:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 18:14 otto@deploy1003: Started deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e]
* 18:13 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e] (duration: 00m 39s)
* 18:13 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e]
* 18:13 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 18:12 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 18:11 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 18:11 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit2002.wikimedia.org with reason: maintenance upgrade
* 18:11 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:11 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 18:10 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit1003.wikimedia.org with reason: maintenance upgrade
* 18:09 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 18:08 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 18:06 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 16:40 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 16:35 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 16:33 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]] synced to the te
* 16:28 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 15:33 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4
* 15:04 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] (duration: 08m 08s)
* 15:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host clouddumps1001.wikimedia.org with OS bookworm
* 14:59 samtar@deploy1003: samtar: Continuing with deployment
* 14:58 samtar@deploy1003: samtar: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:56 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]]
* 14:40 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:07 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 13:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host clouddumps1001.wikimedia.org with OS bookworm
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:55 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:51 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:42 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:41 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:48 klausman@dns1004: END - running authdns-update
* 12:46 klausman@dns1004: START - running authdns-update
* 12:35 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 12:35 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 12:05 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning x4
* 12:01 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2005.codfw.wmnet
* 11:55 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2005.codfw.wmnet
* 11:54 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1144.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:52 cgoubert@dns1004: END - running authdns-update
* 11:49 cgoubert@dns1004: START - running authdns-update
* 11:31 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2004.codfw.wmnet
* 11:25 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2004.codfw.wmnet
* 11:24 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:16 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1200.eqiad.wmnet
* 11:16 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1200.eqiad.wmnet
* 11:04 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1200.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:04 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1199.eqiad.wmnet
* 11:03 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1199.eqiad.wmnet
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:42 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1199.eqiad.wmnet with reason: Upgrading RAID firmware
* 10:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1024.eqiad.wmnet with reason: Cloning x4
* 10:00 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 09:56 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 09:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1024.eqiad.wmnet
* 09:45 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 09:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 09:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:04 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:51 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:46 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning x4
* 08:43 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:34 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] (duration: 09m 56s)
* 08:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:30 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:29 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:28 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:24 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]]
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 08:09 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 07:58 XioNoX: netflow1004:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:56 XioNoX: netflow2005:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 07:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 07:03 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:59 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:43 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:42 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:39 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 06:38 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 06:37 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 05:04 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 05:03 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tool-server' for release 'main' .
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 38s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1345.eqiad.wmnet
* 00:11 dzahn@dns1004: END - running authdns-update
* 00:08 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 00:08 dzahn@dns1004: START - running authdns-update
== 2026-09-09 ==
* 23:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:29 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1345
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 23:28 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1345.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1344.eqiad.wmnet
* 23:22 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] (duration: 11m 15s)
* 23:18 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 23:16 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 23:11 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]]
* 22:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:51 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2004.codfw.wmnet with OS bookworm
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1344
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1344
* 22:37 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1344
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1344
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 22:32 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] (duration: 10m 21s)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1344.eqiad.wmnet
* 22:27 derenrich@deploy1003: derenrich, egardner: Continuing with deployment
* 22:26 derenrich@deploy1003: derenrich, egardner: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:24 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:22 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]]
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:19 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] (duration: 13m 40s)
* 22:16 derenrich@deploy1003: derenrich: Rolling back deployment
* 22:10 derenrich@deploy1003: derenrich: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:05 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]]
* 22:03 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:44 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:40 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:36 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1343
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1343
* 21:27 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1343
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:27 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:23 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:22 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1343
* 21:22 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] (duration: 12m 29s)
* 21:21 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 21:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1343.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1342.eqiad.wmnet
* 21:17 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:14 jforrester@deploy1003: jforrester: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:09 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]]
* 21:08 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:07 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 21:07 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:06 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2004.codfw.wmnet
* 20:59 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:57 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:56 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2004.codfw.wmnet
* 20:53 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:53 bking@cumin2003: END (ERROR) - Cookbook sre.presto.roll-restart-workers (exit_code=97) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:53 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:50 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* {{safesubst:SAL entry|1=20:45 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2}}
* 20:42 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* 20:41 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 20:40 sbassett@deploy1003: aranyap, sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:39 sbassett@deploy1003: aranyap, sbassett: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "Filter}}
* 20:35 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* {{safesubst:SAL entry|1=20:34 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "}}
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:28 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:24 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 20:24 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1342
* 20:23 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 20:23 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1342.eqiad.wmnet
* 20:23 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1342.eqiad.wmnet
* 20:22 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1342.eqiad.wmnet
* 20:15 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1027.eqiad.wmnet with OS bookworm
* 19:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:41 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1027.eqiad.wmnet with OS bookworm
* 19:36 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:28 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:26 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:19 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 19:19 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:06 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 19:05 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 19:03 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 18:59 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 18:23 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:06 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1026.eqiad.wmnet with OS bookworm
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 18:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 17:54 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 17:49 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 17:45 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1026.eqiad.wmnet with OS bookworm
* 17:31 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:29 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 17:27 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 17:12 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 17:04 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 16:46 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] (duration: 09m 28s)
* 16:41 urbanecm@deploy1003: migr, urbanecm: Continuing with deployment
* 16:41 urbanecm@deploy1003: migr, urbanecm: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:36 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]]
* 16:28 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:25 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 15:55 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 15:54 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 15:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] (duration: 09m 43s)
* 15:41 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 15:40 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:36 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]]
* 15:17 urbanecm: Delete all running periodic jobs starting with `growthexperiments-refreshlinkrecommendations-*` (to pick up new configuration; [[phab:T392944|T392944]])
* 15:08 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:06 moritzm: installing grub2 bugfix updates from Bookworm point release
* 15:04 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp6008.drmrs.wmnet
* 15:01 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:42 hnowlan: half concurrency for parsoidCachePrewarm RecordLintJob and refreshLinks in jobqueue, eqiad & codfw
* 14:35 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 14:34 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:32 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 14:20 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:07 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2013.codfw.wmnet with OS trixie
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 14:06 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 13:55 btullis@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'.
* 13:53 btullis@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'.
* 13:43 btullis@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 13:42 btullis@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 13:29 moritzm: pruned obsolete Bullseye image dispatch from the docker registry [[phab:T416452|T416452]]
* 13:28 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:26 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b7-eqiad
* 13:25 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 13:24 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:22 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 13:22 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:17 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a4-eqiad
* 13:17 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] (duration: 10m 15s)
* 13:14 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 13:11 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 13:10 sbisson@deploy1003: sbisson: Continuing with deployment
* 13:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:04 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]]
* 13:02 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on ldap-rw[1001,2001].wikimedia.org with reason: work in progress
* 12:49 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 12:48 btullis@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
* 12:46 btullis@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
* 12:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] (duration: 14m 39s)
* 12:41 ladsgroup@deploy1003: tryvix1509, ladsgroup: Continuing with deployment
* 12:35 ladsgroup@deploy1003: tryvix1509, ladsgroup: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:31 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]]
* 12:16 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 12:16 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 11:53 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] (duration: 21m 58s)
* 11:48 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 11:35 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:31 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]]
* 10:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Repooling db2196
* 10:47 moritzm: pruned obsolete Bullseye images nodejs12-slim/nodejs12-devel/nodejs14-slim/nodejs16-slim from the docker registry [[phab:T416452|T416452]]
* 10:43 moritzm: installing Bird security updates
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1260: Repooling after cloning
* 10:09 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Repooling db2196
* 10:07 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 10:06 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 09:55 moritzm: pruned obsolete Bullseye images openjdk-8-jdk/openjdk-8-jre/openjdk-11-jre/openjdk-11-jdk from the docker registry [[phab:T416452|T416452]]
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1260: Repooling after cloning
* 09:52 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 09:52 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 09:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:28 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:27 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:03 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 09:02 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 09:01 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 16 hosts with reason: upgrade ssw1-a1-eqiad
* 08:58 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 22 hosts with reason: upgrade ssw1-a1-eqiad
* 08:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 08:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 08:49 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 08:49 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 08:48 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 08:48 cmooney@cumin1004: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 08:36 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning sanitarium
* 08:30 brouberol@dns1004: END - running authdns-update
* 08:28 moritzm: pruned obsolete Bullseye image golang1.15 from the docker registry [[phab:T416452|T416452]]
* 08:28 brouberol@dns1004: START - running authdns-update
* 08:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1260: Needs to clone another host from this one
* 08:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1260: Needs to clone another host from this one
* 08:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1260.eqiad.wmnet with reason: Cloning sanitarium
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:00 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 07:40 chlod: UTC morning backport window done
* 07:37 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] (duration: 21m 36s)
* 07:32 chlod@deploy1003: chlod, hamishz: Continuing with deployment
* 07:20 chlod@deploy1003: chlod, hamishz: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 45s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1025.eqiad.wmnet with OS bookworm
== 2026-09-08 ==
* 23:51 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:48 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:39 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1313.eqiad.wmnet
* 23:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:25 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 23:19 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:19 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:15 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:05 Amir1: dropped 57 tables on db1260 ([[phab:T437278|T437278]])
* 23:03 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 23:03 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:02 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 22:57 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1313
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1313
* 22:37 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1313
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:37 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:33 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 22:33 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1313
* 22:32 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 22:32 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1313.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1306.eqiad.wmnet
* 22:27 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6008.drmrs.wmnet with OS trixie
* 22:18 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 22:03 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 22:01 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] (duration: 09m 53s)
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 21:58 Amir1: drop links tables from db2210 ([[phab:T437278|T437278]])
* 21:57 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 21:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:56 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:52 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 21:51 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]]
* 21:51 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:48 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:45 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] (duration: 05m 27s)
* 21:43 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Continuing with deployment
* 21:40 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:39 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]]
* 21:35 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1024.eqiad.wmnet with OS bookworm
* 21:33 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp6008.drmrs.wmnet with OS trixie
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1306
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1306
* 21:30 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1306
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:24 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] (duration: 09m 12s)
* 21:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 21:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1306
* 21:19 reedy@deploy1003: reedy: Continuing with deployment
* 21:19 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:19 reedy@deploy1003: reedy: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1306.eqiad.wmnet
* 21:15 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:15 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]]
* {{safesubst:SAL entry|1=21:10 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out}}
* 21:05 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1024.eqiad.wmnet with OS bookworm
* 21:05 sbassett@deploy1003: sbassett: Continuing with deployment
* 21:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 21:04 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=21:03 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out non-http(s) lice}}
* 20:59 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=20:58 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out n}}
* 20:53 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 20:50 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1305.eqiad.wmnet
* 20:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 20:34 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 20:28 sbassett@deploy1003: sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:27 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-http(s) license}}
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1023.eqiad.wmnet with OS bookworm
* {{safesubst:SAL entry|1=20:23 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-}}
* 20:15 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] (duration: 10m 16s)
* 20:13 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:10 aaron@deploy1003: aaron: Continuing with deployment
* 20:09 aaron@deploy1003: aaron: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 20:05 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:05 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]]
* 20:01 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1023.eqiad.wmnet with OS bookworm
* 19:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1305
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1305
* 19:43 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1305
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:43 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:39 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 19:39 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1305
* 19:38 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 19:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1305.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1275.eqiad.wmnet
* 19:23 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 19:17 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 18:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:53 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1275
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1275
* 18:33 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1275
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1275
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 18:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1275.eqiad.wmnet
* 18:18 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 17:43 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy1003.eqiad.wmnet
* 17:36 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet
* 17:36 cdobbins@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-ntp (exit_code=0) rolling restart_daemons on A:dnsbox
* 17:30 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy1003.eqiad.wmnet
* 17:25 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet
* 17:15 swfrench@deploy1003: Finished scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup (duration: 04m 18s)
* 17:11 Amir1: dropping links tables from db1247 (s4 replica) - ([[phab:T437278|T437278]])
* 17:10 swfrench@deploy1003: Started scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup
* 16:51 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] (duration: 10m 19s)
* 16:46 zabe@deploy1003: zabe: Continuing with deployment
* 16:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]]
* 16:29 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 jhancock@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:12 jhancock@cumin2003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 16:08 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2003.codfw.wmnet
* 15:58 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:55 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2003.codfw.wmnet
* 15:54 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:44 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2002.codfw.wmnet
* 15:19 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2002.codfw.wmnet
* 14:55 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:46 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2001.codfw.wmnet
* 14:44 topranks: shutdown et-1/1/5 on cr1-codfw to shift traffic off ssw1-a1-codfw
* 14:43 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: upgrade ssw1-a1-eqiad
* 14:34 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2001.codfw.wmnet
* 14:33 btullis@cumin1003: END (ERROR) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=97) rolling reboot on A:cephosd-codfw
* 14:30 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet
* 14:28 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet
* 14:22 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet
* 14:13 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --olderThan {{Gerrit|1788220800}} --verbose # [[phab:T437158|T437158]]
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet
* 14:12 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet
* 14:03 topranks: drain traffic from ssw1-a1-codfw before JunOS upgrade [[phab:T426197|T426197]]
* 14:02 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet
* 13:58 cgoubert@deploy1003: helmfile [staging-codfw] DONE helmfile.d/services/mw-debug: apply
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet
* 13:57 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet
* 13:56 cgoubert@deploy1003: helmfile [staging-codfw] START helmfile.d/services/mw-debug: apply
* 13:55 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:54 stran@deploy1003: mwscript-k8s job started: foreachwikiindblist checkuser-suggested-investigations extensions/CheckUser/maintenance/populateSiCaseProperties.php # [[phab:T435066|T435066]]
* 13:52 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:51 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet
* 13:50 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:46 cdobbins@cumin1003: START - Cookbook sre.dns.roll-restart-ntp rolling restart_daemons on A:dnsbox
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet
* 13:46 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet
* 13:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2012.codfw.wmnet with OS bookworm
* 13:44 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] (duration: 34m 00s)
* 13:40 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet
* 13:37 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet
* 13:35 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet
* 13:32 stran@deploy1003: stran: Continuing with deployment
* 13:29 stran@deploy1003: stran: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet
* 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:21 moritzm: installing qemu security updates
* 13:18 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet
* 13:16 ayounsi@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'email' for AS: 139628
* 13:15 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 139628
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet
* 13:13 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet
* 13:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 2519
* 13:11 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 2519
* 13:10 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 14593
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]]
* 13:09 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:08 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 14593
* 13:06 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet
* 13:06 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2012.codfw.wmnet with OS bookworm
* 13:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2012.codfw.wmnet
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2012.codfw.wmnet
* 13:05 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:04 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet
* 13:01 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet
* 12:58 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 34655
* 12:58 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 34655
* 12:56 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2012.codfw.wmnet
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'clear' for AS: 35320
* 12:55 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'clear' for AS: 35320
* 12:55 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-codfw
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-codfw
* 12:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2011.codfw.wmnet with OS bookworm
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-by27-esams
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-by27-esams
* 12:54 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-esams
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-esams
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-esams
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-esams
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-eqsin
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-eqsin
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-ulsfo
* 12:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f3-eqiad
* 12:50 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f3-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-c8-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-c8-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:44 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:43 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet
* 12:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-eqiad
* 12:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f1-eqiad
* 12:41 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f1-eqiad
* 12:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-d5-eqiad
* 12:40 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-d5-eqiad
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet
* 12:38 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet
* 12:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:32 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet
* 12:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad
* 12:11 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2012.codfw.wmnet
* 12:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2011.codfw.wmnet with OS bookworm
* 12:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2011.codfw.wmnet
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2011.codfw.wmnet
* 12:09 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:06 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] (duration: 09m 54s)
* 12:01 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2011.codfw.wmnet
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Continuing with deployment
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2010.codfw.wmnet with OS bookworm
* 11:56 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]]
* 11:46 marostegui@dns1004: END - running authdns-update
* 11:44 marostegui@dns1004: START - running authdns-update
* 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:40 Amir1: dropping unneeded tables from x4 - db1260 ([[phab:T437278|T437278]])
* 11:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:23 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2011.codfw.wmnet
* 11:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2010.codfw.wmnet with OS bookworm
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2010.codfw.wmnet
* 11:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2010.codfw.wmnet
* 11:20 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2010.codfw.wmnet
* 11:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2009.codfw.wmnet with OS bookworm
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:43 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] (duration: 10m 57s)
* 10:39 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:38 samtar@deploy1003: samtar: Continuing with deployment
* 10:37 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:37 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:36 samtar@deploy1003: samtar: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:34 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:32 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]]
* 10:30 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2009.codfw.wmnet with OS bookworm
* 10:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2009.codfw.wmnet
* 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:18 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2009.codfw.wmnet
* 10:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2008.codfw.wmnet with OS bookworm
* 10:07 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:05 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:01 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:45 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] (duration: 13m 15s)
* 09:45 ayounsi@dns1004: END - running authdns-update
* 09:44 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:43 ayounsi@dns1004: START - running authdns-update
* 09:39 zabe@deploy1003: zabe: Continuing with deployment
* 09:37 zabe@deploy1003: zabe: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2008.codfw.wmnet with OS bookworm
* 09:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:32 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]]
* 09:32 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2008.codfw.wmnet
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2008.codfw.wmnet
* 09:31 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:29 XioNoX: remove GRE tunnels eqiad-drmrs eqdfw-ulsfo
* 09:23 moritzm: installing rsync security updates
* 09:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2007.codfw.wmnet with OS bookworm
* 09:22 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2008.codfw.wmnet
* 09:11 marostegui@cumin1003: dbctl commit (dc=all): 'Make x4 and s4 RW again [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96393 and previous config saved to /var/cache/conftool/dbconfig/20260908-091121-marostegui.json
* 09:07 marostegui@cumin1003: dbctl commit (dc=all): 'Remove old s4 masters from x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96392 and previous config saved to /var/cache/conftool/dbconfig/20260908-090749-marostegui.json
* 09:05 marostegui@cumin1003: dbctl commit (dc=all): 'Set x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96391 and previous config saved to /var/cache/conftool/dbconfig/20260908-090517-marostegui.json
* 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Set s4 commons to read-only for maintenance [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96389 and previous config saved to /var/cache/conftool/dbconfig/20260908-090228-marostegui.json
* 09:02 marostegui: Starting x4 split from s4, RO time on commons needed [[phab:T404715|T404715]]
* 09:00 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 08:58 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2008.codfw.wmnet
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:43 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 32 hosts with reason: x4 split
* 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2007.codfw.wmnet with OS bookworm
* 08:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:37 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:37 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2007.codfw.wmnet
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2007.codfw.wmnet
* 08:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:34 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:29 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:26 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2007.codfw.wmnet
* 08:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2006.codfw.wmnet with OS bookworm
* 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:50 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2007.codfw.wmnet
* 07:43 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2006.codfw.wmnet with OS bookworm
* 07:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2006.codfw.wmnet
* 07:37 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:30 denisse: Add grafana-plugins 0.15 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 07:29 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:27 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 07:27 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:27 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:22 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 07:22 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 07:18 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:14 jmm@dns1004: END - running authdns-update
* 07:13 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96388 and previous config saved to /var/cache/conftool/dbconfig/20260908-071308-marostegui.json
* 07:12 jmm@dns1004: START - running authdns-update
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96387 and previous config saved to /var/cache/conftool/dbconfig/20260908-071216-marostegui.json
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96386 and previous config saved to /var/cache/conftool/dbconfig/20260908-071159-marostegui.json
* 05:07 denisse: Add grafana-plugins 0.10 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.16 (duration: 02m 27s)
* 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]] (duration: 36m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 41s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-07 ==
* 21:52 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] (duration: 11m 00s)
* 21:47 zabe@deploy1003: zabe: Continuing with deployment
* 21:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]]
* 21:37 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] (duration: 09m 34s)
* 21:33 zabe@deploy1003: zabe: Continuing with deployment
* 21:32 zabe@deploy1003: zabe: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:28 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]]
* 21:03 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] (duration: 10m 27s)
* 20:58 zabe@deploy1003: zabe: Continuing with deployment
* 20:57 zabe@deploy1003: zabe: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:52 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]]
* 20:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set weight of db1261 to zero in s4 ([[phab:T437108|T437108]])', diff saved to https://phabricator.wikimedia.org/P96385 and previous config saved to /var/cache/conftool/dbconfig/20260907-203804-ladsgroup.json
* 20:23 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] (duration: 09m 28s)
* 20:19 zabe@deploy1003: zabe: Continuing with deployment
* 20:18 zabe@deploy1003: zabe: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:14 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]]
* 20:12 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] (duration: 10m 06s)
* 20:07 zabe@deploy1003: zabe, daimona: Continuing with deployment
* 20:06 zabe@deploy1003: zabe, daimona: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:02 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]]
* 19:59 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] (duration: 11m 27s)
* 19:55 zabe@deploy1003: zabe: Continuing with deployment
* 19:52 zabe@deploy1003: zabe: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:48 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]]
* 19:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] (duration: 11m 40s)
* 19:26 zabe@deploy1003: zabe: Continuing with deployment
* 19:23 zabe@deploy1003: zabe: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:19 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]]
* 19:07 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] (duration: 14m 17s)
* 19:00 zabe@deploy1003: zabe: Continuing with deployment
* 18:57 zabe@deploy1003: zabe: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:53 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]]
* 18:33 zabe@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 18:32 zabe@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 18:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] (duration: 09m 12s)
* 18:27 zabe@deploy1003: zabe: Continuing with deployment
* 18:26 zabe@deploy1003: zabe: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:22 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]]
* 16:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2005.codfw.wmnet with OS bookworm
* 16:01 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] (duration: 10m 22s)
* 15:59 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 15:57 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 15:56 zabe@deploy1003: zabe: Continuing with deployment
* 15:55 zabe@deploy1003: zabe: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:51 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]]
* 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:47 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 15:46 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 15:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:44 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 15:44 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 15:27 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2005.codfw.wmnet with OS bookworm
* 15:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2005.codfw.wmnet
* 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:11 moritzm: installing rsync security updates
* 15:04 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1341.eqiad.wmnet
* 15:03 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2004.codfw.wmnet with OS bookworm
* 15:00 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:58 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 14:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:40 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 14:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 14:35 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:32 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:29 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:25 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 14:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1228: Repooling db1228 into s4
* 14:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2004.codfw.wmnet with OS bookworm
* 14:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Repooling after cloning
* 14:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:19 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2004.codfw.wmnet
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2004.codfw.wmnet
* 14:18 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:14 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1074.eqiad.wmnet
* 14:14 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:13 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:13 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* {{safesubst:SAL entry|1=14:11 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mul}}
* 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2004.codfw.wmnet
* 14:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1074.eqiad.wmnet
* 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1073.eqiad.wmnet
* 14:07 krinkle@deploy1003: krinkle: Continuing with deployment
* {{safesubst:SAL entry|1=14:04 krinkle@deploy1003: krinkle: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with multiple properties}}
* 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1073.eqiad.wmnet
* 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1072.eqiad.wmnet
* 14:01 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1341
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1341
* 14:01 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 14:00 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1341
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* 13:59 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* {{safesubst:SAL entry|1=13:59 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mult}}
* 13:59 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2004.codfw.wmnet
* 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2003.codfw.wmnet with OS bookworm
* 13:55 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 13:55 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1072.eqiad.wmnet
* 13:55 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudvirt1067.eqiad.wmnet
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1341
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 13:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1341.eqiad.wmnet
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2008.wikimedia.org
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2008.wikimedia.org with OS trixie
* 13:39 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1067.eqiad.wmnet
* 13:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1066.eqiad.wmnet
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:37 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1228: Repooling db1228 into s4
* 13:36 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Repooling after cloning
* 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1066.eqiad.wmnet
* 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1065.eqiad.wmnet
* 13:28 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:27 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1065.eqiad.wmnet
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:25 moritzm: installing openssh security updates
* 13:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:24 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1340.eqiad.wmnet
* 13:23 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] (duration: 10m 06s)
* 13:17 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2003.codfw.wmnet with OS bookworm
* 13:17 stran@deploy1003: stran: Continuing with deployment
* 13:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:16 stran@deploy1003: stran: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:16 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:15 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 13:12 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]]
* 13:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2003.codfw.wmnet
* 13:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2003.codfw.wmnet
* 13:03 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2008.wikimedia.org with OS trixie
* 13:03 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:02 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:02 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:01 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:00 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2003.codfw.wmnet
* 12:58 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:54 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 12:54 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2008.wikimedia.org
* 12:47 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2003.codfw.wmnet
* 12:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2002.codfw.wmnet with OS bookworm
* 12:29 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 12:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2007.wikimedia.org
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2007.wikimedia.org with OS trixie
* 12:14 elukey: moved most of the Docker Registry's prefixes to a new internal S3 backend. For any docker pull failure that worked in the past, please ping me or drop a note in [[phab:T435499|T435499]] or contact the oncall SREs
* 12:07 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2002.codfw.wmnet with OS bookworm
* 12:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:02 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2002.codfw.wmnet
* 11:54 jmm@dns1004: END - running authdns-update
* 11:52 jmm@dns1004: START - running authdns-update
* 11:46 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2002.codfw.wmnet
* 11:46 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2007.wikimedia.org with OS trixie
* 11:46 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:46 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:41 moritzm: installing bash updates from bookworm point release
* 11:39 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:39 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2007.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts ldap-replica1006.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:35 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:32 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2001.codfw.wmnet with OS bookworm
* 11:28 jmm@cumin1004: START - Cookbook sre.hosts.decommission for hosts ldap-replica1006.wikimedia.org
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 11:18 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] (duration: 14m 08s)
* 11:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 11:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet
* 11:11 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:11 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:10 zabe@deploy1003: zabe: Continuing with deployment
* 11:10 zabe@deploy1003: zabe: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 11:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet
* 11:07 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet
* 11:07 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet
* 11:06 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:06 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:04 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]]
* 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 10:58 btullis@deploy1003: Finished scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]] (duration: 35m 20s)
* 10:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:53 jmm@dns1004: END - running authdns-update
* 10:51 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:50 jmm@dns1004: START - running authdns-update
* 10:47 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=frwiki # [[phab:T436659|T436659]]
* 10:40 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=hrwiki # [[phab:T436659|T436659]]
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:39 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 10:37 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 10:33 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 10:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2001.codfw.wmnet with OS bookworm
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1340
* 10:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1340.eqiad.wmnet
* 10:26 btullis@deploy1003: Started scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]]
* 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2001.codfw.wmnet
* 10:18 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2001.codfw.wmnet
* 10:13 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw
* 10:12 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:10 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:09 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:08 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:07 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:07 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:02 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:00 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:59 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:57 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/services/miscweb: apply
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/services/miscweb: apply
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:52 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:52 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] (duration: 10m 11s)
* 09:51 aokoth@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 09:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 09:49 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 09:49 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 09:48 aokoth@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 09:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:45 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:44 aokoth@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 09:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]]
* 09:41 aokoth@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 09:38 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Repooling after switchover
* 09:23 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw
* 09:15 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] (duration: 20m 12s)
* 09:12 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 139009
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1005.wikimedia.org
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1005.wikimedia.org with OS trixie
* 09:10 moritzm: rebuild software RAID following disk replacement [[phab:T437036|T437036]]
* 09:10 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 139009
* 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1022.eqiad.wmnet with OS bookworm
* 09:08 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 09:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:03 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet
* 09:02 moritzm: installing giflib security updates
* 09:01 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:59 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 08:57 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 08:56 jmm@cumin1004: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet
* 08:56 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:54 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]]
* 08:52 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:52 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:49 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 08:48 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:42 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:41 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:40 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:40 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96355 and previous config saved to /var/cache/conftool/dbconfig/20260907-083904-marostegui.json
* 08:38 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1005.wikimedia.org with OS trixie
* 08:38 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2157 to s5 primary [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96354 and previous config saved to /var/cache/conftool/dbconfig/20260907-083825-marostegui.json
* 08:38 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 marostegui: Starting s5 codfw failover from db2213 to db2157 - [[phab:T437188|T437188]]
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2157 with weight 0 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96353 and previous config saved to /var/cache/conftool/dbconfig/20260907-083448-marostegui.json
* 08:34 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s5 [[phab:T437188|T437188]]
* 08:28 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:28 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1005.wikimedia.org
* 08:22 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:20 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1022.eqiad.wmnet with OS bookworm
* 08:03 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:02 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 08:02 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:00 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 07:57 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 07:56 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1241.eqiad.wmnet with reason: Cloning
* 07:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1241: Cloning
* 07:55 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1241: Cloning
* 07:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Cloning
* 07:54 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Cloning
* 07:51 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 07:50 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 07:47 kartik@deploy1003: Finished scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] (duration: 41m 51s)
* 07:46 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 07:45 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 07:34 kartik@deploy1003: abi, kartik: Continuing with deployment
* 07:23 kartik@deploy1003: abi, kartik: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:05 kartik@deploy1003: Started scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]]
* 06:14 moritzm: installing Chromium security updates
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 08m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-06 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 25s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-05 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 26s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-04 ==
* 22:07 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 21:48 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:42 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:30 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:42 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:40 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 19:27 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:19 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:13 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:12 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host sretest2013
* 19:11 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host sretest2013
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:11 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:07 jhancock@cumin1003: START - Cookbook sre.dns.netbox
* 18:18 inflatador: bking@clouddumps100[12] `systemctl reset-failed` to quash alerts until https://w.wiki/UBje . The systemd timer should try again tomorrow
* 17:27 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b8-eqiad
* 17:27 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b8-eqiad
* 16:37 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:33 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-b7-eqiad
* 16:33 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 16:05 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b6-eqiad
* 16:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b6-eqiad
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1339.eqiad.wmnet
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 15:50 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:49 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:47 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:43 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 15:10 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b5-eqiad
* 15:09 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b5-eqiad
* 14:46 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1045.eqiad.wmnet
* 14:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1003.eqiad.wmnet with OS trixie
* 14:38 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 14:37 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b4-eqiad
* 14:37 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b4-eqiad
* 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339
* 14:32 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339
* 14:31 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 14:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet
* 14:26 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:24 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:21 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b3-eqiad
* 14:21 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b3-eqiad
* 14:17 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:17 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:04 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1003.eqiad.wmnet with OS trixie
* 13:54 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b2-eqiad
* 13:53 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b2-eqiad
* 13:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1002.eqiad.wmnet with OS trixie
* 13:18 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a4-eqiad
* 13:12 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:12 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b1-eqiad
* 13:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b1-eqiad
* 12:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1002.eqiad.wmnet with OS trixie
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow3004.esams.wmnet with OS trixie
* 12:33 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:28 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 12:15 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 12:11 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a7-eqiad
* 12:11 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a7-eqiad
* 12:01 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow3004.esams.wmnet with OS trixie
* 11:47 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a6-eqiad
* 11:47 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a6-eqiad
* 11:36 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki2003.codfw.wmnet
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki2003.codfw.wmnet with OS trixie
* 11:19 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:13 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a5-eqiad
* 11:06 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a5-eqiad
* 10:52 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki2003.codfw.wmnet with OS trixie
* 10:50 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:50 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:44 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 10:44 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki2003.codfw.wmnet
* 10:29 btullis@deploy1003: Finished scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli (duration: 41m 14s)
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 10:00 marostegui@cumin1003: Removing db1182 from zarcillo [[phab:T434869|T434869]]
* 10:00 marostegui@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db1182.eqiad.wmnet
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:57 btullis@deploy1003: Started scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli
* 09:53 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:53 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.decommission (exit_code=1)
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1182.eqiad.wmnet
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:50 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:46 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:45 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1182 from dbctl [[phab:T434869|T434869]]', diff saved to https://phabricator.wikimedia.org/P96346 and previous config saved to /var/cache/conftool/dbconfig/20260904-094527-marostegui.json
* 09:41 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki1003.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki1003.eqiad.wmnet with OS trixie
* 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1182: Decommissioning
* 09:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1182: Decommissioning
* 09:23 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 09:17 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow5003.eqsin.wmnet with OS trixie
* 09:02 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki1003.eqiad.wmnet with OS trixie
* 09:00 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 09:00 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:55 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:55 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki1003.eqiad.wmnet
* 08:54 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:48 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:45 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:40 btullis@deploy1003: Finished scap sync-world: Trying again for [[phab:T436913|T436913]] (duration: 34m 26s)
* 08:35 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2003.codfw.wmnet
* 08:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw2001.wikimedia.org with OS trixie
* 08:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4 days, 0:00:00 on db2196.codfw.wmnet with reason: Host crashed
* 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2003.codfw.wmnet
* 08:24 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:20 elukey@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:12 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:08 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:07 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 08:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:04 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 07:57 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow5003.eqsin.wmnet with OS trixie
* 07:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:55 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw2001.wikimedia.org with OS trixie
* 07:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:51 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:50 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie
* 07:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2196: down
* 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2196: down
* 06:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:52 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:40 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:28 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:21 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 06:20 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 39s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 01:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:02 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:59 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:46 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 00:44 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:39 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:23 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
== 2026-09-03 ==
* 21:46 tsev@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T435363|T435363]]
* 21:03 eevans@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:55 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] (duration: 12m 24s)
* 20:52 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:52 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1021.eqiad.wmnet with OS bookworm
* 20:50 arlolra@deploy1003: arlolra, tsev: Continuing with deployment
* 20:46 arlolra@deploy1003: arlolra, tsev: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:44 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1047.eqiad.wmnet
* 20:42 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]]
* 20:41 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a3-eqiad
* 20:40 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a3-eqiad
* 20:39 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] (duration: 10m 23s)
* 20:34 arlolra@deploy1003: arlolra, jgiannelos: Continuing with deployment
* 20:33 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1047.eqiad.wmnet
* 20:33 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:32 arlolra@deploy1003: arlolra, jgiannelos: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:31 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1046.eqiad.wmnet
* 20:28 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]]
* 20:23 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] (duration: 13m 41s)
* 20:20 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1046.eqiad.wmnet
* 20:16 catrope@deploy1003: catrope: Continuing with deployment
* 20:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1021.eqiad.wmnet with OS bookworm
* 20:13 catrope@deploy1003: catrope: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1021.eqiad.wmnet
* 20:09 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]]
* 20:00 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1021.eqiad.wmnet
* 19:49 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 19:19 swfrench@deploy1003: Finished scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]] (duration: 02m 59s)
* 19:16 swfrench@deploy1003: Started scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]]
* 19:01 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:01 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 18:57 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:57 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:34 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:20 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 17:55 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:55 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:53 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:52 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:46 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 17:46 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 17:45 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 17:44 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 17:44 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:40 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:39 ryankemper: [WDQS] Service looks healthy again, CPU load and thread count have dropped considerably over the last hour
* 17:39 ryankemper: [[phab:T421642|T421642]] [WDQS] requestctl changes: `2026-09-03 16:23-17:33` UTC: added hard-deny pair `cache-text/wdqs_futile_sparql_sep_2026_deny(+_bots)`; extended pattern `ua/wdqs_heavy_sparql_bots_2026` and added default-scope twin `wdqs_heavy_sparql_bots_jul_2026_ratelimit_default`; added ipblock `abuse/wdqs_sparql_scanners_sep_2026` + throttle `wdqs_sparql_scanners_sep_2026_ratelimit` (needed manual `requestctl update-provenance-map`)
* 17:37 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:37 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:31 andrew@cumin2003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudcephosd1045.eqiad.wmnet
* 17:30 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:30 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:28 dancy@deploy1003: Installation of scap version "4.289.0" completed for 3 hosts
* 17:26 dancy@deploy1003: Installing scap version "4.289.0" for 3 host(s)
* 17:24 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a2-eqiad
* 17:24 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:22 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:22 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a2-eqiad
* 17:14 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 17:10 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:10 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:09 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 17:01 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:59 andrew@cumin2003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 16:58 dancy: Running scap clean-images on deploy1003
* 16:52 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:50 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:39 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:14 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs2021.codfw.wmnet,service=wdqs-main
* 16:14 ryankemper: [[phab:T430880|T430880]] Stumbled across `wdqs2021` listed as inactive, looks like it was never fully re-pooled after a data xfer. Pooled.
* 16:12 btullis@deploy1003: Finished deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2] (duration: 00m 38s)
* 16:12 btullis@deploy1003: Started deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2]
* 16:12 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:07 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs101[1-4].eqiad.wmnet
* 16:03 btullis@deploy1003: Started scap sync-world: Rebuilding to pick up new version of dump scripts in mediawiki-cli for [[phab:T436913|T436913]]
* 16:01 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]] (duration: 09m 29s)
* 15:51 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=urldownloader[12]00[56].wikimedia.org [reason: depooling urldownloader trixie nodes]
* 15:51 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]]
* 15:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:29 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:24 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:24 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:18 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:15 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:09 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:05 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 15:04 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:02 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 15:00 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader
* 14:58 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw
* 14:53 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-ulsfo
* 14:49 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 14:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2002.codfw.wmnet
* 14:45 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:39 sukhe: sudo cumin "A:cp-text" "run-puppet-agent --enable 'merging CR 1334855'": [[phab:T425441|T425441]]
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2002.codfw.wmnet
* 14:32 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:30 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:29 sukhe: sudo cumin "A:cp-text" "disable-puppet 'merging CR 1334855'" [[phab:T425441|T425441]]
* 14:27 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-ulsfo
* 14:22 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:21 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:19 arnaudb@dns1006: END - running authdns-update
* 14:18 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1074
* 14:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1074
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-codfw
* 14:17 arnaudb@dns1006: START - running authdns-update
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:16 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:16 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1073
* 14:15 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1073
* 14:12 ayounsi@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host netflow2004.codfw.wmnet with OS trixie
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:10 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:08 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:08 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:07 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] (duration: 09m 36s)
* 14:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:05 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:03 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 14:03 samtar@deploy1003: btullis, samtar: Continuing with deployment
* 14:02 samtar@deploy1003: btullis, samtar: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:01 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:58 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]]
* 13:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:56 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:55 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:52 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqiad and A:durum
* 13:52 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-codfw
* 13:51 moritzm: installing sqlite3 security updates
* 13:51 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:51 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqiad and A:durum
* 13:49 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-codfw and A:durum
* 13:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:47 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-codfw and A:durum
* 13:47 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-esams
* 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:44 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:43 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:43 ayounsi@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:42 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] (duration: 13m 50s)
* 13:41 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:40 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1072
* 13:40 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:39 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-esams and A:durum
* 13:38 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1072
* 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:38 samtar@deploy1003: samtar, thiemowmde: Continuing with deployment
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-esams and A:durum
* 13:37 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:33 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:33 samtar@deploy1003: samtar, thiemowmde: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:32 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqsin and A:durum
* 13:31 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqsin and A:durum
* 13:28 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]]
* 13:28 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 13:27 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:24 ayounsi@cumin1004: START - Cookbook sre.hosts.reimage for host netflow2004.codfw.wmnet with OS trixie
* 13:24 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:22 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 13:22 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-esams
* 13:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:15 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:15 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 moritzm: installing bash updates from trixie point release
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 13:14 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1067
* 13:13 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1067
* 13:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2901: Test
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:09 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:09 moritzm: installing libxslt bugfix updates from Trixie point release
* 13:08 jelto@dns1004: END - running authdns-update
* 13:06 jelto@dns1004: START - running authdns-update
* 13:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 13:04 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:04 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 13:00 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:59 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1066
* 12:59 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1066
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:58 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:54 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 12:53 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2901: Test
* 12:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2901: Test
* 12:52 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'sync'.
* 12:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 12:50 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2901: Test
* 12:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:50 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 12:42 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:41 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-magru and A:durum
* 12:38 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-magru and A:durum
* 12:34 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 12:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1282: Pooling db1282 into s6
* 12:31 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 12:25 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:21 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:19 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:15 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_ulsfo
* 12:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1020.eqiad.wmnet with OS bookworm
* 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: db2207 repool
* 12:07 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_ulsfo
* 12:04 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 11:58 kart_: cxserver: Use urldownloader LVS endpoint ([[phab:T429175|T429175]])
* 11:57 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply
* 11:56 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply
* 11:56 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_magru
* 11:55 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply
* 11:55 moritzm: installing rsync security updates
* 11:55 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply
* 11:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:52 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply
* 11:51 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply
* 11:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1282: Pooling db1282 into s6
* 11:45 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1282 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96328 and previous config saved to /var/cache/conftool/dbconfig/20260903-114526-marostegui.json
* 11:43 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqiad
* 11:35 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqiad
* 11:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1020.eqiad.wmnet with OS bookworm
* 11:26 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 11:24 cgoubert@deploy1003: Finished scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter (duration: 12m 01s)
* 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: db2207 repool
* 11:22 cgoubert@deploy1003: cgoubert: Continuing with deployment
* 11:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqsin
* 11:17 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_esams
* 11:15 cgoubert@deploy1003: cgoubert: mediawiki: enable forward of fatal metrics to statsd exporter synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:14 cgoubert@deploy1003: Started scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter
* 11:10 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_esams
* 11:09 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 11:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1019.eqiad.wmnet with OS bookworm
* 11:01 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_drmrs
* 10:59 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_codfw
* 10:52 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_codfw
* 10:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:41 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 10:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 10:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow2005.codfw.wmnet
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2005.codfw.wmnet with OS trixie
* 10:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:17 btullis@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync
* 10:17 btullis@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync
* 10:16 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1019.eqiad.wmnet with OS bookworm
* 10:15 btullis@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: sync
* 10:15 btullis@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: sync
* 10:12 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s8
* 10:11 marostegui: Move s8 sanitarium from db1167 to db1281 [[phab:T434778|T434778]]
* 10:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 10:03 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 09:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:55 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2003.codfw.wmnet with OS trixie
* 09:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:43 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow2005.codfw.wmnet with OS trixie
* 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:42 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1018.eqiad.wmnet with OS bookworm
* 09:41 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_ulsfo
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:39 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2003.codfw.wmnet with reason: host reimage
* 09:39 hnowlan: fixed currently oncall pane in klaxon
* 09:38 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:38 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434751|T434751]]
* 09:37 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:37 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow2005.codfw.wmnet
* 09:35 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434775|T434775]]
* 09:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s7
* 09:34 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:33 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_ulsfo
* 09:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqiad
* 09:30 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] (duration: 09m 30s)
* 09:27 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:25 zabe@deploy1003: zabe: Continuing with deployment
* 09:25 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqiad
* 09:25 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 09:25 zabe@deploy1003: zabe: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1174 from dbctl [[phab:T436904|T436904]]', diff saved to https://phabricator.wikimedia.org/P96323 and previous config saved to /var/cache/conftool/dbconfig/20260903-092448-marostegui.json
* 09:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:21 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]]
* 09:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqsin
* 09:17 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:15 topranks: put traffic on Lumen codfw<->eqiad link as it is stable [[phab:T435810|T435810]]
* 09:14 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_esams
* 09:09 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:06 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_esams
* 09:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 09:03 marostegui: Move s6 sanitarium from db1165 to db1279 [[phab:T434775|T434775]]
* 09:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1018.eqiad.wmnet with OS bookworm
* 08:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s6
* 08:57 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_drmrs
* 08:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_codfw
* 08:55 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw
* 08:49 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_codfw
* 08:49 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:45 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 08:45 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:42 marostegui: Move s5 sanitarium from db1161 to db1275 [[phab:T434776|T434776]]
* 08:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 24 hosts with reason: Changing sanitarium master in s5
* 08:38 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_magru
* 08:37 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:37 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server (duration: 01m 21s)
* 08:37 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:36 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server
* 08:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:33 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server (duration: 01m 28s)
* 08:32 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server
* 08:27 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 08:11 moritzm: uploaded wmf-laptop 1.0.7 to apt.wikimedia.org
* 08:03 marostegui: Move s2 sanitarium from db1156 to db1271 [[phab:T434287|T434287]]
* 07:59 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2001.codfw.wmnet
* 07:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 23 hosts with reason: Changing sanitarium master in s2
* 07:49 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2001.codfw.wmnet
* 07:39 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:29 chlod: UTC morning backport window done
* 07:27 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] (duration: 11m 54s)
* 07:22 chlod@deploy1003: chlod, tryvix1509: Continuing with deployment
* 07:22 XioNoX: push pfw policies - [[phab:T436729|T436729]]
* 07:20 chlod@deploy1003: chlod, tryvix1509: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]]
* 07:15 marostegui: Power off db1228 for maintenance
* 07:13 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1228.eqiad.wmnet with reason: Onsite maintenance
* 07:01 arnaudb@dns1006: END - running authdns-update
* 06:58 arnaudb@dns1006: START - running authdns-update
* 06:54 jmm@cumin2003: END (PASS) - Cookbook sre.wdqs.restart-nginx-envoy (exit_code=0) rolling restart_daemons on A:wcqs-public
* 06:52 jmm@cumin2003: START - Cookbook sre.wdqs.restart-nginx-envoy rolling restart_daemons on A:wcqs-public
* 06:46 moritzm: installing libxml2 security updates
* 06:27 hashar: Upgrading CI Jenkins on contint1003 # [[phab:T436812|T436812]]
* 06:11 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet
* 06:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts2002.codfw.wmnet
* 06:00 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet
* 05:56 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts2002.codfw.wmnet
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 48s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:16 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1017.eqiad.wmnet with OS bookworm
== 2026-09-02 ==
* 23:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] (duration: 10m 21s)
* 23:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:50 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 23:49 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]]
* 23:38 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:38 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host aqs1017.eqiad.wmnet with OS bookworm
* 22:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] (duration: 11m 03s)
* 22:32 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 22:30 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:26 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]]
* 22:05 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] (duration: 14m 14s)
* 21:59 krinkle@deploy1003: krinkle: Continuing with deployment
* 21:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:50 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]]
* 21:44 inflatador: bking@apt1002 sudo -E private_reprepro --ignore=wrongdistribution -C matomo_plugins include bookworm-wikimedia-private matomo-plugin-customreports_5.5.0-1_amd64.changes [[phab:T431608|T431608]]
* 21:40 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] (duration: 09m 48s)
* 21:35 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:34 jforrester@deploy1003: jforrester: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:34 inflatador: bking@apt1002 sudo -E reprepro -C main include bookworm-wikimedia matomo-plugin-marketingcampaignsreporting_5.2.2-3_amd64.changes [[phab:T431608|T431608]]
* 21:30 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]]
* 21:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 21:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:03 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 20:50 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 20:27 dancy@deploy1003: Finished scap sync-world: testing (duration: 09m 21s)
* 20:18 dancy@deploy1003: Started scap sync-world: testing
* 20:18 dancy@deploy1003: Installation of scap version "4.288.0" completed for 3 hosts
* 20:16 dancy@deploy1003: Installing scap version "4.288.0" for 3 host(s)
* 19:57 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] (duration: 64m 27s)
* 19:55 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 18:57 jforrester@deploy1003: jforrester: Continuing with deployment
* 18:57 jforrester@deploy1003: jforrester: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:55 swfrench-wmf: deleted pods coredns-85b4f68d95-pk5sn coredns-85b4f68d95-22ddb coredns-85b4f68d95-49k5p in eqiad due to intermittent upstream resolution health check failures correlated with high DNS resolution latency
* 18:53 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]]
* 18:31 sukhe@dns1004: END - running authdns-update
* 18:28 sukhe@dns1004: START - running authdns-update
* 18:26 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:26 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:17 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:16 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 18:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:14 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:14 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqiad
* 18:14 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:49 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqiad
* 17:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:45 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqsin
* 17:38 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:38 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:20 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqsin
* 17:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 17:00 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 16:59 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:48 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 16:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:41 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 16:41 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 16:41 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:38 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:34 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 16:29 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_drmrs
* 16:24 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 16:14 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_drmrs
* 16:11 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-unlock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 16:10 root@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]] (duration: 48m 09s)
* 16:10 root@deploy1003: Forcefully removing global lock: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 16:10 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-unlock-scap for datacenter switchover from codfw to eqiad
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:59 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:58 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters for datacenter switchover from codfw to eqiad
* 15:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 15:58 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-restore-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-restore-ttl for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-start-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 root@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-start-maintenance for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-jobrunner: sync
* 15:56 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-jobrunner: sync
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.07-set-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period ends at: 2026-09-02 15:56:13.434320
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.07-set-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.03-set-db-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.03-set-db-readonly for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.02-set-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period starts at: 2026-09-02 15:53:47.690918
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.02-set-readonly for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.01-stop-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.01-stop-maintenance for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-reduce-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:47 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-reduce-ttl for datacenter switchover from codfw to eqiad
* 15:46 slyngshede@cumin1003: END (ERROR) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=97) for datacenter switchover from codfw to eqiad
* 15:45 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 15:44 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 15:42 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service
* 15:39 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:38 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:31 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:28 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:27 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 15:27 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from codfw to eqiad
* 15:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-lock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:22 root@deploy1003: Locking from deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-lock-scap for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks for datacenter switchover from codfw to eqiad
* 15:17 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service
* 15:15 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_magru
* 15:11 moritzm: import jenkins 2.568.3 to thirdparty/jenkins for trixie-wikimedia
* 14:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 14:44 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_magru
* 14:32 moritzm: installing pdns-recursor security updates
* 14:27 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:27 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:20 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:15 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 14:12 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 14:09 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:09 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:09 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:08 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:06 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:44 moritzm: bounce tcpircbot-logmsgbot/tcpircbot-logmsgbot_cloud on alert1002 to allow cumin1004 [[phab:T427897|T427897]]
* 13:36 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] (duration: 09m 52s)
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Continuing with deployment
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:26 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]]
* 13:25 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 13:24 moritzm: installing wireshark security updates
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 13:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:04 moritzm: import librsvg 2.60.0+dfsg-1+wmf13u1 to component/thumbor for trixie-wikimedia [[phab:T436505|T436505]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 12:44 atsuko@dns1004: END - running authdns-update
* 12:41 atsuko@dns1004: START - running authdns-update
* 12:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:35 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] (duration: 12m 50s)
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:24 dreamyjazz@deploy1003: dreamyjazz, btullis: Continuing with deployment
* 12:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:22 dreamyjazz@deploy1003: dreamyjazz, btullis: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:20 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:17 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]]
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 11:31 marostegui@cumin1003: Removing db1172 from zarcillo [[phab:T436763|T436763]]
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1172.eqiad.wmnet
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:30 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:26 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:26 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:25 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 11:25 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:24 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:20 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1172.eqiad.wmnet
* 11:20 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 11:12 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:11 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:10 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:09 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:05 slyngshede@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.*
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 11:05 slyngshede@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 11:03 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply
* 11:03 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
* 10:52 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:49 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:48 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Pool back db1242
* 10:45 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow4003.ulsfo.wmnet with OS trixie
* 10:31 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1172 from dbctl [[phab:T436763|T436763]]', diff saved to https://phabricator.wikimedia.org/P96318 and previous config saved to /var/cache/conftool/dbconfig/20260902-103152-marostegui.json
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 blake@deploy1003: Finished scap sync-world: non-build deployment for [[phab:T417800|T417800]] (duration: 05m 37s)
* 10:06 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:05 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:04 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:03 blake@deploy1003: Started scap sync-world: non-build deployment for [[phab:T417800|T417800]]
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 10:00 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:00 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:59 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:58 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:57 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:56 jmm@dns1004: END - running authdns-update
* 09:56 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:55 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:54 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:53 jmm@dns1004: START - running authdns-update
* 09:51 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:47 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1228 to dbctl [[phab:T435892|T435892]]', diff saved to https://phabricator.wikimedia.org/P96313 and previous config saved to /var/cache/conftool/dbconfig/20260902-094713-marostegui.json
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:42 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow4003.ulsfo.wmnet with OS trixie
* 09:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow7002.magru.wmnet with OS trixie
* 09:31 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:30 moritzm: installing openjdk-21 security updates
* 09:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:22 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:17 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:17 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:16 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:10 moritzm: installing openjdk-8 security updates
* 09:09 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:08 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:46 tappof: bump space for prometheus k8s-dse in eqiad
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:39 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow7002.magru.wmnet with OS trixie
* 08:36 moritzm: installing libgraphite2 security updates
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:23 Msz2001: UTC morning backport window done
* 08:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] (duration: 14m 36s)
* 08:19 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:19 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet
* 08:14 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:14 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet
* 08:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet
* 08:08 fabfur@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cp5022.eqsin.wmnet with reason: investigating
* 08:07 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]]
* 08:07 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp5022.*
* 08:07 fabfur: depooling and silencing cp5022 ([[phab:T414411|T414411]])
* 08:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet
* 08:03 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:03 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* {{safesubst:SAL entry|1=08:03 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)}}
* 07:49 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 07:49 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]] synced to the
* 07:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1004.eqiad.wmnet
* 07:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1004.eqiad.wmnet
* 07:30 jmm@dns1004: END - running authdns-update
* {{safesubst:SAL entry|1=07:27 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]}}
* 07:27 jmm@dns1004: START - running authdns-update
* 07:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] (duration: 16m 04s)
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1242: Cloning db1228
* 07:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1242: Cloning db1228
* 07:18 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1228,1242].eqiad.wmnet with reason: db1242 needs to clone db1228
* 07:17 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Continuing with deployment
* 07:12 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1228.eqiad.wmnet with OS trixie
* 07:10 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be veri
* 07:06 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]]
* 06:43 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: build: Updating npm dependencies (duration: 00m 13s)
* 06:43 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: build: Updating npm dependencies
* 06:39 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:32 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:23 slyngshede@dns1004: END - running authdns-update
* 06:21 marostegui: Drop cu* tables from s3 bswiktionary [[phab:T435965|T435965]]
* 06:20 slyngshede@dns1004: START - running authdns-update
* 06:18 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1228.eqiad.wmnet with OS trixie
* 06:13 XioNoX: re-enable magru cr1/asw1-b3 link - [[phab:T436675|T436675]]
* 05:06 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] (duration: 04m 42s)
* 05:04 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:03 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:01 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 05:01 tstarling@deploy1003: Scap cancelled without rolling back.
* 04:53 tstarling@deploy1003: tstarling: Continuing with deployment
* 04:29 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 04:25 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 43s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-01 ==
* 21:59 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] (duration: 18m 05s)
* 21:52 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:47 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]]
* 21:38 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] (duration: 23m 55s)
* 21:28 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:20 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:14 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]]
* 20:47 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1016.eqiad.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:24 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:11 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 20:11 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:57 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:57 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:56 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:54 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:52 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:47 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:45 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:42 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:41 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:40 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:39 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:32 jhancock@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudcephosd1055
* 19:32 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:30 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.* [reason: update IP addrs]
* 19:30 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 19:30 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:23 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:22 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:16 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:13 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1056
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1056
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1055
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:12 jclark@cumin1003: END (ERROR) - Cookbook sre.network.configure-switch-interfaces (exit_code=97) for host cloudcephosd1054
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1054
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:11 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:06 jclark@cumin1003: START - Cookbook sre.dns.netbox
* 19:06 sukhe@dns1004: END - running authdns-update
* 19:03 sukhe@dns1004: START - running authdns-update
* 18:14 dancy@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 17:04 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 17:00 dancy@deploy1003: Finished scap sync-world: testing (duration: 08m 07s)
* 16:52 dancy@deploy1003: Started scap sync-world: testing
* 16:48 dancy@deploy1003: sync-world aborted: testing (duration: 00m 05s)
* 16:48 dancy@deploy1003: Started scap sync-world: testing
* 16:47 dancy@deploy1003: Installation of scap version "4.287.0" completed for 156 hosts
* 16:42 dancy@deploy1003: Installing scap version "4.287.0" for 156 host(s)
* 16:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 16:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 15:51 moritzm: installing mesa security updates
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts phab1004.eqiad.wmnet
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:27 aokoth@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:20 aokoth@cumin1003: START - Cookbook sre.dns.netbox
* 15:14 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595] (duration: 05m 32s)
* 15:14 aokoth@cumin1003: START - Cookbook sre.hosts.decommission for hosts phab1004.eqiad.wmnet
* 15:11 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on phab1004.eqiad.wmnet with reason: Decom
* 15:09 joal@deploy1003: Started deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595]
* 15:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config: apply
* 15:08 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config: apply
* 14:55 hashar: Restarted Jenkins on releases1003
* 14:51 hashar: Restarted CI Jenkins on contint1003
* 14:48 hashar: Restarting Gerrit primary on gerrit2003
* 14:45 hashar: Restarted Gerrit on gerrit1003 and gerrit2002
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 14:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 14:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 14:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:24 moritzm: installing curl security updates
* 14:24 jmm@dns1004: END - running authdns-update
* 14:23 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser (duration: 00m 14s)
* 14:23 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:21 jmm@dns1004: START - running authdns-update
* 14:21 jmm@dns1004: END - running authdns-update
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 14:19 jmm@dns1004: START - running authdns-update
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet
* 14:14 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:14 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:13 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595] (duration: 07m 26s)
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Test
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Test
* 14:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet
* 14:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Test
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Test
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 14:05 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595]
* 14:05 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] (duration: 37m 37s)
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 hashar: Removed openjdk-17 packages from contint1002/contint2002 following relocation of CI Jenkins to contint1003/contint2003 # [[phab:T418521|T418521]]
* 14:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:02 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:02 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:00 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595] (duration: 00m 45s)
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 13:59 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595]
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:58 ladsgroup@dns1004: END - running authdns-update
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:57 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 ladsgroup@dns1004: START - running authdns-update
* 13:56 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:56 ladsgroup@dns1004: END - running authdns-update
* 13:55 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:53 ladsgroup@dns1004: START - running authdns-update
* 13:50 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:49 kharlan@deploy1003: kharlan: Continuing with deployment
* 13:48 kharlan@deploy1003: kharlan: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2004.wikimedia.org
* 13:42 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:41 jmm@dns1004: END - running authdns-update
* 13:40 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2004.wikimedia.org
* 13:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2003.wikimedia.org
* 13:38 jmm@dns1004: START - running authdns-update
* 13:34 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2003.wikimedia.org
* 13:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1004.wikimedia.org
* 13:26 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]]
* 13:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:25 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:24 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1004.wikimedia.org
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:23 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] (duration: 20m 24s)
* 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1003.wikimedia.org
* 13:20 fnegri@deploy1003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply
* 13:19 fnegri@deploy1003: helmfile [eqiad] START helmfile.d/services/toolhub: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:16 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1003.wikimedia.org
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:15 moritzm: bump urldownloader[12]00[34] to 8G RAM [[phab:T429175|T429175]]
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:15 fnegri@deploy1003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 fnegri@deploy1003: helmfile [codfw] START helmfile.d/services/toolhub: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:11 fnegri@deploy1003: helmfile [staging] DONE helmfile.d/services/toolhub: apply
* 13:11 aude@deploy1003: aude: Continuing with deployment
* 13:10 fnegri@deploy1003: helmfile [staging] START helmfile.d/services/toolhub: apply
* 13:07 aude@deploy1003: aude: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2004.wikimedia.org with OS bookworm
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow1004.eqiad.wmnet
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1004.eqiad.wmnet with OS trixie
* 12:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:47 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:41 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] (duration: 16m 25s)
* 12:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 12:34 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 12:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 12:33 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 12:31 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:29 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:25 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]]
* 12:22 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1004.eqiad.wmnet with OS trixie
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:20 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 12:20 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 12:16 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow1004.eqiad.wmnet
* 12:15 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2004.wikimedia.org with OS bookworm
* 12:14 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 12:07 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2003.wikimedia.org with OS bookworm
* 11:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:43 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:32 jmm@cumin2003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling restart_daemons on A:kafka-test-eqiad
* 11:26 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2003.wikimedia.org with OS bookworm
* 11:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1004.wikimedia.org with OS bookworm
* 11:12 moritzm: installing openjdk-21 security updates
* 11:12 jmm@cumin2003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling restart_daemons on A:kafka-test-eqiad
* 10:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:53 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:49 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2902: Pool back db2902
* 10:45 moritzm: installing Python 3.11 security updates
* 10:37 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1004.wikimedia.org with OS bookworm
* 10:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5022.eqsin.wmnet with OS trixie
* 10:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp5022.eqsin.wmnet on all recursors
* 10:36 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache cp5022.eqsin.wmnet on all recursors
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:30 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:04 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:04 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2902: Pool back db2902
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:03 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2902: test
* 10:03 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 10:01 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 09:58 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:50 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/echoserver: apply
* 09:49 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/echoserver: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Change db1176 and db2230's weight, test-s4 masters, to 0 to mimic the rest of production [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96292 and previous config saved to /var/cache/conftool/dbconfig/20260901-092444-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96291 and previous config saved to /var/cache/conftool/dbconfig/20260901-090233-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96290 and previous config saved to /var/cache/conftool/dbconfig/20260901-090158-marostegui.json
* 09:01 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96289 and previous config saved to /var/cache/conftool/dbconfig/20260901-090121-marostegui.json
* 09:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:56 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:50 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1003.wikimedia.org with OS bookworm
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1074.eqiad.wmnet
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:46 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:43 marostegui@cumin1003: dbctl commit (dc=all): 'Test repool db2902', diff saved to https://phabricator.wikimedia.org/P96288 and previous config saved to /var/cache/conftool/dbconfig/20260901-084317-marostegui.json
* 08:42 marostegui@cumin1003: dbctl commit (dc=all): 'Test depool db2902', diff saved to https://phabricator.wikimedia.org/P96287 and previous config saved to /var/cache/conftool/dbconfig/20260901-084249-marostegui.json
* 08:39 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:37 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2902: test
* 08:36 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96285 and previous config saved to /var/cache/conftool/dbconfig/20260901-083557-marostegui.json
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96284 and previous config saved to /var/cache/conftool/dbconfig/20260901-083527-marostegui.json
* 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96283 and previous config saved to /var/cache/conftool/dbconfig/20260901-083432-marostegui.json
* 08:32 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1074.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1073.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:27 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:24 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 08:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022']
* 08:16 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1073.eqiad.wmnet
* 08:15 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1003.wikimedia.org with OS bookworm
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1072.eqiad.wmnet
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:14 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:10 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:05 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1072.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1067.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:00 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:56 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:54 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022']
* 07:53 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 07:52 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie
* 07:52 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 07:50 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1067.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1066.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:47 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 07:41 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:34 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1066.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1065.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:32 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:27 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:23 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1065.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1075.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:15 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 06:49 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:45 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1075.eqiad.wmnet
* 06:29 moritzm: installing Java 17 security updates
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.15 (duration: 02m 25s)
* 03:50 denisse@deploy1003: Finished deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2 (duration: 00m 19s)
* 03:50 denisse@deploy1003: Started deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2
* 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]] (duration: 37m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:30 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 00:29 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
== Other archives ==
See [[Server Admin Log/Archives]].
<noinclude>
[[Category:SAL]]
[[Category:Operations]]
</noinclude>
5ab8876k69rvm3ik04x284ly3lefhxc
2458706
2458704
2026-09-19T20:08:36Z
JrandWP
37706
archive Sept 1-18
2458706
wikitext
text/x-wiki
== 2026-09-19 ==
* 16:55 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:11 urbanecm: Attach SHB@commonswiki to the SUL account manually ([[phab:T438591|T438591]], see [[phab:T438591|T438591]]#12341750 for what I did exactly)
* 04:08 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 04:08 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 04:08 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 04:07 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 36s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== Other archives ==
See [[Server Admin Log/Archives]].
<noinclude>
[[Category:SAL]]
[[Category:Operations]]
</noinclude>
7ty3rj6qv15k2bsc4d2q3srnai675tk
Map of database maintenance
0
449160
2458708
2458685
2026-09-20T00:00:56Z
Dexbot
30554
Bot: Updating the report
2458708
wikitext
text/x-wiki
{{/Header}}
== Today (2026-09-20) ==
== Yesterday (2026-09-19) ==
== Last seven days ==
{| class="wikitable"
|+ eqiad
|-
! Section !! Work
|-
| x1 || [[phab:T436496|Pre switchover: Database preparations (T436496)]] (marostegui)
|-
|}
[[Category:MariaDB]]
qg7q3msjatgfltzeadslvck0gdbijee
Server Admin Log/Archive 110
0
460698
2458705
2026-09-19T20:08:12Z
JrandWP
37706
Archive Sept 1-18
2458705
wikitext
text/x-wiki
== 2026-09-18 ==
* 22:41 rzl@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=sessionstore,name=eqiad
* 17:08 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343065{{!}}mw.DesktopArticleTarget: if source education is enabled suppress welcome (T434249)]] (duration: 09m 26s)
* 17:05 Dreamy_Jazz: Created `securepoll_log` on `nlwiki` main DB cluster for [[phab:T434045|T434045]]
* 17:04 kemayo@deploy1003: kemayo: Continuing with deployment
* 17:03 kemayo@deploy1003: kemayo: Backport for [[gerrit:1343065{{!}}mw.DesktopArticleTarget: if source education is enabled suppress welcome (T434249)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:59 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1343065{{!}}mw.DesktopArticleTarget: if source education is enabled suppress welcome (T434249)]]
* 16:49 oblivian@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343127{{!}}ResourceLoader: hotfix for current logspam over the weekend (T438387)]] (duration: 11m 36s)
* 16:42 oblivian@deploy1003: oblivian: Continuing with deployment
* 16:42 oblivian@deploy1003: oblivian: Backport for [[gerrit:1343127{{!}}ResourceLoader: hotfix for current logspam over the weekend (T438387)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:37 oblivian@deploy1003: Started scap sync-world: Backport for [[gerrit:1343127{{!}}ResourceLoader: hotfix for current logspam over the weekend (T438387)]]
* 16:09 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ml-serve1016.eqiad.wmnet with OS trixie
* 14:49 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie
* 13:37 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host registry1004.eqiad.wmnet with OS trixie
* 13:23 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on registry1004.eqiad.wmnet with reason: host reimage
* 13:18 elukey@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on registry1004.eqiad.wmnet with reason: host reimage
* 13:04 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host registry1004.eqiad.wmnet with OS trixie
* 12:14 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:13 jclark@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:30 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 09:30 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 09:22 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on an-redacteddb1001.eqiad.wmnet with reason: cloning
* 09:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 09:19 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 09:18 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 21 hosts with reason: cloning db1270
* 09:18 marostegui: clone db1270:x4 from db1155:x4 lag will appear on x4
* 09:09 brouberol@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 09:08 brouberol@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 08:23 brouberol@dns1004: END - running authdns-update
* 08:21 brouberol@dns1004: START - running authdns-update
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 57s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-17 ==
* 21:04 tsev@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T438395|T438395]]
* 20:55 tsev@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T438395|T438395]]
* 20:47 jhuneidi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342774{{!}}Worklist Promotion test kitchen - Enable flag in production (T434513)]], [[gerrit:1342781{{!}}Exclude returntoapp query from app interception on iOS (T438395)]], [[gerrit:1342798{{!}}Revert "Update wikimania wordmark for 2026"]] (duration: 35m 59s)
* 20:35 jhuneidi@deploy1003: robertsky, jhuneidi, cmelo, tsev: Continuing with deployment
* 20:31 jhuneidi@deploy1003: robertsky, jhuneidi, cmelo, tsev: Backport for [[gerrit:1342774{{!}}Worklist Promotion test kitchen - Enable flag in production (T434513)]], [[gerrit:1342781{{!}}Exclude returntoapp query from app interception on iOS (T438395)]], [[gerrit:1342798{{!}}Revert "Update wikimania wordmark for 2026"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:11 jhuneidi@deploy1003: Started scap sync-world: Backport for [[gerrit:1342774{{!}}Worklist Promotion test kitchen - Enable flag in production (T434513)]], [[gerrit:1342781{{!}}Exclude returntoapp query from app interception on iOS (T438395)]], [[gerrit:1342798{{!}}Revert "Update wikimania wordmark for 2026"]]
* 19:15 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1100.eqiad.wmnet with OS trixie
* 19:15 vriley@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1004"
* 19:10 vriley@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1004"
* 18:52 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1100.eqiad.wmnet with reason: host reimage
* 18:48 vriley@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1100.eqiad.wmnet with reason: host reimage
* 18:32 vriley@cumin1004: START - Cookbook sre.hosts.reimage for host ms-be1100.eqiad.wmnet with OS trixie
* 18:20 urbanecm: Deploy a security fix for [[phab:T438389|T438389]]
* 17:47 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be1100.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:38 vriley@cumin1004: START - Cookbook sre.hosts.provision for host ms-be1100.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:38 vriley@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1100.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:37 vriley@cumin1004: START - Cookbook sre.hosts.provision for host ms-be1100.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:23 vriley@cumin1004: START - Cookbook sre.hosts.reimage for host ms-be1100.eqiad.wmnet with OS trixie
* 16:52 aokoth@deploy1003: Finished deploy [phabricator/deployment@c386249]: Deploy Phab (duration: 00m 12s)
* 16:52 aokoth@deploy1003: Started deploy [phabricator/deployment@c386249]: Deploy Phab
* 16:42 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1099.eqiad.wmnet with OS trixie
* 16:42 vriley@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1004"
* 16:42 vriley@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1004"
* 16:35 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be1100.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:31 aokoth@deploy1003: Finished deploy [phabricator/deployment@c386249]: Deploy Phab (duration: 00m 19s)
* 16:31 aokoth@deploy1003: Started deploy [phabricator/deployment@c386249]: Deploy Phab
* 16:21 vriley@cumin1004: START - Cookbook sre.hosts.provision for host ms-be1100.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:20 vriley@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be1100
* 16:20 vriley@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host ms-be1100
* 16:19 vriley@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:19 vriley@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [ms-be1100] - vriley@cumin1004"
* 16:19 vriley@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [ms-be1100] - vriley@cumin1004"
* 16:15 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 16:15 vriley@cumin1004: START - Cookbook sre.dns.netbox
* 16:15 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 16:14 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 16:14 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 16:12 moritzm: installing libapache-mod-auth-oidc security updates
* 16:12 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1099.eqiad.wmnet with reason: host reimage
* 16:11 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 16:11 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 16:08 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/editcheck-headless: apply
* 16:07 rzl@deploy1003: helmfile [staging] START helmfile.d/services/editcheck-headless: apply
* 16:06 vriley@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1099.eqiad.wmnet with reason: host reimage
* 16:01 moritzm: installing aom security updates
* 16:01 btullis@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ceph-admin2001.codfw.wmnet
* 16:01 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ceph-admin2001.codfw.wmnet with OS bookworm
* 15:55 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 15:55 rzl@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 15:54 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 15:54 rzl@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 15:54 rzl@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 15:53 rzl@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 15:53 rzl@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 15:52 rzl@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 15:51 vriley@cumin1004: START - Cookbook sre.hosts.reimage for host ms-be1099.eqiad.wmnet with OS trixie
* 15:48 moritzm: installing libde265 security updates
* 15:44 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ceph-admin2001.codfw.wmnet with reason: host reimage
* 15:40 btullis@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ceph-admin1001.eqiad.wmnet
* 15:40 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ceph-admin1001.eqiad.wmnet with OS bookworm
* 15:39 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir4003.*
* 15:37 btullis@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ceph-admin2001.codfw.wmnet with reason: host reimage
* 15:26 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir4003.ulsfo.wmnet with OS trixie
* 15:25 aokoth@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host phab2003.codfw.wmnet with OS trixie
* 15:23 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ceph-admin1001.eqiad.wmnet with reason: host reimage
* 15:19 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host registry1005.eqiad.wmnet with OS trixie
* 15:17 btullis@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ceph-admin1001.eqiad.wmnet with reason: host reimage
* 15:16 btullis@cumin1004: START - Cookbook sre.hosts.reimage for host ceph-admin2001.codfw.wmnet with OS bookworm
* 15:16 btullis@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ceph-admin2001.codfw.wmnet - btullis@cumin1004"
* 15:16 btullis@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ceph-admin2001.codfw.wmnet - btullis@cumin1004"
* 15:15 btullis@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ceph-admin2001.codfw.wmnet on all recursors
* 15:15 btullis@cumin1004: START - Cookbook sre.dns.wipe-cache ceph-admin2001.codfw.wmnet on all recursors
* 15:15 btullis@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:15 btullis@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ceph-admin2001.codfw.wmnet - btullis@cumin1004"
* 15:15 btullis@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ceph-admin2001.codfw.wmnet - btullis@cumin1004"
* 15:08 aokoth@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on phab2003.codfw.wmnet with reason: host reimage
* 15:06 Msz2001: Deployed private code changes to Suggestedinvestigations
* 15:05 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir4003.ulsfo.wmnet with reason: host reimage
* 15:05 aokoth@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on phab2003.codfw.wmnet with reason: host reimage
* 15:04 btullis@cumin1004: START - Cookbook sre.hosts.reimage for host ceph-admin1001.eqiad.wmnet with OS bookworm
* 15:04 btullis@cumin1004: START - Cookbook sre.dns.netbox
* 15:04 btullis@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ceph-admin1001.eqiad.wmnet - btullis@cumin1004"
* 15:04 btullis@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ceph-admin1001.eqiad.wmnet - btullis@cumin1004"
* 15:04 btullis@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ceph-admin1001.eqiad.wmnet on all recursors
* 15:04 btullis@cumin1004: START - Cookbook sre.dns.wipe-cache ceph-admin1001.eqiad.wmnet on all recursors
* 15:04 btullis@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:04 btullis@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ceph-admin1001.eqiad.wmnet - btullis@cumin1004"
* 15:04 btullis@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ceph-admin1001.eqiad.wmnet - btullis@cumin1004"
* 15:02 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on registry1005.eqiad.wmnet with reason: host reimage
* 15:01 btullis@cumin1004: START - Cookbook sre.ganeti.makevm for new host ceph-admin2001.codfw.wmnet
* 15:00 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342265{{!}}JsonSchemaBuilder: Cache the root schema in the process (T437588)]], [[gerrit:1342264{{!}}JsonSchemaBuilder: Cache the root schema in the process (T437588)]], [[gerrit:1342684{{!}}SI: Preserve the username filter when switching queues (T438308)]] (duration: 12m 34s)
* 15:00 btullis@cumin1004: START - Cookbook sre.dns.netbox
* 15:00 btullis@cumin1004: START - Cookbook sre.ganeti.makevm for new host ceph-admin1001.eqiad.wmnet
* 14:58 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir4003.ulsfo.wmnet with reason: host reimage
* 14:58 elukey@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on registry1005.eqiad.wmnet with reason: host reimage
* 14:55 urbanecm@deploy1003: mszwarc, urbanecm: Continuing with deployment
* 14:51 urbanecm@deploy1003: mszwarc, urbanecm: Backport for [[gerrit:1342265{{!}}JsonSchemaBuilder: Cache the root schema in the process (T437588)]], [[gerrit:1342264{{!}}JsonSchemaBuilder: Cache the root schema in the process (T437588)]], [[gerrit:1342684{{!}}SI: Preserve the username filter when switching queues (T438308)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:51 aokoth@cumin1004: START - Cookbook sre.hosts.reimage for host phab2003.codfw.wmnet with OS trixie
* 14:50 aokoth@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on phab2003.codfw.wmnet with reason: Reimage
* 14:47 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1342265{{!}}JsonSchemaBuilder: Cache the root schema in the process (T437588)]], [[gerrit:1342264{{!}}JsonSchemaBuilder: Cache the root schema in the process (T437588)]], [[gerrit:1342684{{!}}SI: Preserve the username filter when switching queues (T438308)]]
* 14:42 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1331830{{!}}Keep Balinese Palm Leaf variants enabled on wikisource (T436398)]], [[gerrit:1340216{{!}}Turn on variant conversion for PageAssessments (T328012)]], [[gerrit:1341949{{!}}Parsoid Read Views: Enable on 61 wikiquote wikis (T437917)]] (duration: 15m 31s)
* 14:39 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host registry1005.eqiad.wmnet with OS trixie
* 14:35 cscott@deploy1003: ssastry, cscott: Continuing with deployment
* 14:33 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir4003.ulsfo.wmnet with OS trixie
* 14:32 cscott@deploy1003: ssastry, cscott: Backport for [[gerrit:1331830{{!}}Keep Balinese Palm Leaf variants enabled on wikisource (T436398)]], [[gerrit:1340216{{!}}Turn on variant conversion for PageAssessments (T328012)]], [[gerrit:1341949{{!}}Parsoid Read Views: Enable on 61 wikiquote wikis (T437917)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:26 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1331830{{!}}Keep Balinese Palm Leaf variants enabled on wikisource (T436398)]], [[gerrit:1340216{{!}}Turn on variant conversion for PageAssessments (T328012)]], [[gerrit:1341949{{!}}Parsoid Read Views: Enable on 61 wikiquote wikis (T437917)]]
* 14:20 caro@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342678{{!}}enwiki desktop VE: add education popup for switching to source editor (T434249)]], [[gerrit:1342362{{!}}Make VE the default editor on enwiki desktop (T436574)]] (duration: 33m 52s)
* 14:07 caro@deploy1003: caro: Continuing with deployment
* 14:06 caro@deploy1003: caro: Backport for [[gerrit:1342678{{!}}enwiki desktop VE: add education popup for switching to source editor (T434249)]], [[gerrit:1342362{{!}}Make VE the default editor on enwiki desktop (T436574)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:46 caro@deploy1003: Started scap sync-world: Backport for [[gerrit:1342678{{!}}enwiki desktop VE: add education popup for switching to source editor (T434249)]], [[gerrit:1342362{{!}}Make VE the default editor on enwiki desktop (T436574)]]
* 13:37 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 13:34 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:33 elukey@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:31 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:30 elukey@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:29 Emperor: apus - radosgw-admin quota set --quota-scope=user --uid=docker-registry --max-size=5T [[phab:T438339|T438339]]
* 13:24 jclark@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1016
* 13:24 jclark@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1016
* 13:11 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:11 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:09 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:09 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:55 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:53 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 12:42 jclark@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:40 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:40 jclark@cumin1003: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:35 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 12:19 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342626{{!}}MathMathML: Simplify Mathoid fallback/a11y class logic (T436026)]], [[gerrit:1342627{{!}}ext.math.mathjax: Implement mwe-math-mathml-a11y for client-side MathJax (T436026)]] (duration: 13m 04s)
* 12:14 krinkle@deploy1003: krinkle: Continuing with deployment
* 12:10 krinkle@deploy1003: krinkle: Backport for [[gerrit:1342626{{!}}MathMathML: Simplify Mathoid fallback/a11y class logic (T436026)]], [[gerrit:1342627{{!}}ext.math.mathjax: Implement mwe-math-mathml-a11y for client-side MathJax (T436026)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:05 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1342626{{!}}MathMathML: Simplify Mathoid fallback/a11y class logic (T436026)]], [[gerrit:1342627{{!}}ext.math.mathjax: Implement mwe-math-mathml-a11y for client-side MathJax (T436026)]]
* 10:37 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 10:28 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 10:10 blake@deploy1003: Finished scap sync-world: cleanup for [[phab:T417800|T417800]] (duration: 03m 57s)
* 10:07 blake@deploy1003: Started scap sync-world: cleanup for [[phab:T417800|T417800]]
* 09:52 marostegui@cumin1004: dbctl commit (dc=all): 'Fix weights [[phab:T436496|T436496]]', diff saved to https://phabricator.wikimedia.org/P96466 and previous config saved to /var/cache/conftool/dbconfig/20260917-095235-marostegui.json
* 09:51 marostegui@cumin1004: dbctl commit (dc=all): 'Fix weights [[phab:T436496|T436496]]', diff saved to https://phabricator.wikimedia.org/P96465 and previous config saved to /var/cache/conftool/dbconfig/20260917-095131-marostegui.json
* 09:41 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:41 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 09:41 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:40 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 09:40 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:40 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 09:35 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 09:35 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 08:37 moritzm: pruned obsolete Bullseye image prometheus-nutcracker-exporter from the docker registry [[phab:T416452|T416452]]
* 08:34 XioNoX: Manually install gnmic 0.49.0 on netflow2005 - [[phab:T438291|T438291]]
* 08:28 brouberol@dns1004: END - running authdns-update
* 08:26 brouberol@dns1004: START - running authdns-update
* 08:13 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.20 refs [[phab:T430839|T430839]]
* 08:10 moritzm: imported nodejs_26.8.2-1nodesource1 to thirdparty/node26 for trixie-wikimedia [[phab:T437510|T437510]]
* 08:07 Amir1: dropped links tables from db2206 ([[phab:T437278|T437278]])
* 08:03 Amir1: dropped links tables from db2219 ([[phab:T437278|T437278]])
* 08:01 Amir1: dropped links tables from db2236 ([[phab:T437278|T437278]])
* 07:59 Amir1: dropped non-links tables from db1262 ([[phab:T437278|T437278]])
* 07:57 Amir1: dropped non-links tables from db2245 ([[phab:T437278|T437278]])
* 07:52 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@8eaca67] (releasing): [[phab:T438205|T438205]] to prod host (duration: 00m 44s)
* 07:52 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@8eaca67] (releasing): [[phab:T438205|T438205]] to prod host
* 07:49 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@8eaca67] (releasing): [[phab:T438205|T438205]] to backup host (duration: 00m 47s)
* 07:48 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@8eaca67] (releasing): [[phab:T438205|T438205]] to backup host
* 07:25 XioNoX: Manually install gnmic 0.49.0 on netflow1004 - [[phab:T438291|T438291]]
* 07:24 mlitn@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342398{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342396{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342395{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342545{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342546{{!}}Localisation updates from https://translatewiki.net.]],
* 07:19 mlitn@deploy1003: mlitn, jdlrobson: Continuing with deployment
* {{safesubst:SAL entry|1=07:18 mlitn@deploy1003: mlitn, jdlrobson: Backport for [[gerrit:1342398{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342396{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342395{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342545{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342546{{!}}Localisation updates from https://translatewiki.net.]], [[gerri}}
* 07:11 mlitn@deploy1003: Started scap sync-world: Backport for [[gerrit:1342398{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342396{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342395{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342545{{!}}Localisation updates from https://translatewiki.net.]], [[gerrit:1342546{{!}}Localisation updates from https://translatewiki.net.]],
* 07:10 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Jmoore111 out of all services on: 2444 hosts
* 06:07 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 05:53 marostegui@cumin1004: END (FAIL) - Cookbook sre.mysql.decommission (exit_code=99)
* 05:53 marostegui@cumin1004: Removing db1180 from zarcillo [[phab:T437222|T437222]]
* 05:53 marostegui@cumin1004: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1180.eqiad.wmnet
* 05:53 marostegui@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 05:53 marostegui@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1180.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1004"
* 05:53 marostegui@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1180.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1004"
* 05:49 marostegui@cumin1004: START - Cookbook sre.dns.netbox
* 05:44 marostegui@cumin1004: START - Cookbook sre.hosts.decommission for hosts db1180.eqiad.wmnet
* 05:43 marostegui@cumin1004: START - Cookbook sre.mysql.decommission
* 04:26 aokoth@cumin1004: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet
* 04:24 aokoth@cumin1004: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet
== 2026-09-16 ==
* 23:10 rzl: rzl@deploy1003 Finished scap sync-world: Backport for [[gerrit:1342091{{!}}Repool poolcounter[1007,2006] (T435163)]] (duration: 11m 09s)
* 22:50 rzl@deploy1003: Started scap sync-world: Backport for [[gerrit:1342091{{!}}Repool poolcounter[1007,2006] (T435163)]]
* 22:47 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342383{{!}}DonorIdentification: Confirm before unlinking donor status in preferences (T436698)]], [[gerrit:1342385{{!}}Make learn more link to new window (T438252)]] (duration: 35m 21s)
* 22:35 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 22:33 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1342383{{!}}DonorIdentification: Confirm before unlinking donor status in preferences (T436698)]], [[gerrit:1342385{{!}}Make learn more link to new window (T438252)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:12 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1342383{{!}}DonorIdentification: Confirm before unlinking donor status in preferences (T436698)]], [[gerrit:1342385{{!}}Make learn more link to new window (T438252)]]
* 22:10 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2006.codfw.wmnet
* 22:06 rzl@cumin2003: START - Cookbook sre.hosts.reboot-single for host poolcounter2006.codfw.wmnet
* 22:06 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1007.eqiad.wmnet
* 22:02 rzl@cumin2003: START - Cookbook sre.hosts.reboot-single for host poolcounter1007.eqiad.wmnet
* 21:56 rzl@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342090{{!}}Repool poolcounter[1006,2005]; depool poolcounter[1007,2006] for reboot (T435163)]] (duration: 09m 39s)
* 21:52 rzl@deploy1003: rzl: Continuing with deployment
* 21:51 rzl@deploy1003: rzl: Backport for [[gerrit:1342090{{!}}Repool poolcounter[1006,2005]; depool poolcounter[1007,2006] for reboot (T435163)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:47 rzl@deploy1003: Started scap sync-world: Backport for [[gerrit:1342090{{!}}Repool poolcounter[1006,2005]; depool poolcounter[1007,2006] for reboot (T435163)]]
* 21:46 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:43 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2005.codfw.wmnet
* 21:42 vriley@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1099.eqiad.wmnet with OS trixie
* 21:41 vriley@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1098.eqiad.wmnet with OS trixie
* 21:41 vriley@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin2003"
* 21:40 vriley@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin2003"
* 21:40 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:40 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:39 rzl@cumin2003: START - Cookbook sre.hosts.reboot-single for host poolcounter2005.codfw.wmnet
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1006.eqiad.wmnet
* 21:38 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:38 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:37 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:35 rzl@cumin2003: START - Cookbook sre.hosts.reboot-single for host poolcounter1006.eqiad.wmnet
* 21:32 rzl@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342089{{!}}Depool poolcounter[1006,2005] for reboot (T435163)]] (duration: 13m 53s)
* 21:26 rzl@deploy1003: rzl: Continuing with deployment
* 21:25 rzl@deploy1003: rzl: Backport for [[gerrit:1342089{{!}}Depool poolcounter[1006,2005] for reboot (T435163)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:23 vriley@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1098.eqiad.wmnet with reason: host reimage
* 21:18 rzl@deploy1003: Started scap sync-world: Backport for [[gerrit:1342089{{!}}Depool poolcounter[1006,2005] for reboot (T435163)]]
* 21:17 vriley@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1098.eqiad.wmnet with reason: host reimage
* 21:10 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:09 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:09 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:09 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:08 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:02 vriley@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS trixie
* 20:49 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342339{{!}}Reapply "Tell VisualEditor about the app web edit tags", modified]] (duration: 35m 55s)
* 20:37 kemayo@deploy1003: cklimas, kemayo: Continuing with deployment
* 20:33 kemayo@deploy1003: cklimas, kemayo: Backport for [[gerrit:1342339{{!}}Reapply "Tell VisualEditor about the app web edit tags", modified]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:13 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1342339{{!}}Reapply "Tell VisualEditor about the app web edit tags", modified]]
* 19:20 dwisehaupt@dns1005: END - running authdns-update
* 19:18 dwisehaupt@dns1005: START - running authdns-update
* 19:06 vriley@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:57 dwisehaupt@dns1005: END - running authdns-update
* 18:55 vriley@cumin2003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:55 dwisehaupt@dns1005: START - running authdns-update
* 18:44 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 18:42 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 18:37 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 18:35 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 18:26 robh@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:22 robh@cumin2003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:21 dzahn@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 18:21 robh@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:21 robh@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:20 dzahn@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 18:20 dzahn@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 18:18 dzahn@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 18:18 mutante: k8s/miscweb: admin_ng deploy: creating namespace for attribution.wikimedia.org [[phab:T437635|T437635]]
* 18:17 dzahn@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 18:17 dzahn@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 18:17 dzahn@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 18:16 dzahn@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 18:13 cdanis@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "fix known-client creation - cdanis@cumin1003"
* 18:13 cdanis@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: fix known-client creation - cdanis@cumin1003
* 18:12 cdanis@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: fix known-client creation - cdanis@cumin1003
* 18:12 cdanis@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "fix known-client creation - cdanis@cumin1003"
* 18:04 vriley@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:51 vriley@cumin2003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:46 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 17:46 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 17:45 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:45 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:44 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:44 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:35 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1006.eqiad.wmnet with OS trixie
* 17:35 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1003"
* 17:30 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:29 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1003"
* 17:27 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:27 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:23 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:23 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:20 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 17:20 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 17:14 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:13 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1006.eqiad.wmnet with reason: host reimage
* 17:10 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:10 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:10 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:10 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:09 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:07 jclark@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1006.eqiad.wmnet with reason: host reimage
* 17:07 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 17:05 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:05 vriley@cumin1003: START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:52 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie
* 16:46 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir7003.magru.wmnet
* 16:44 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir7003
* 16:19 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie
* 15:42 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply
* 15:42 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply
* 15:42 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 15:42 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir7003.magru.wmnet with OS trixie
* 15:41 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 15:36 jnuche@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342279{{!}}Use parser output value instead of status (T438154)]] (duration: 33m 21s)
* 15:35 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply
* 15:35 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply
* 15:24 jnuche@deploy1003: jnuche, jforrester: Continuing with deployment
* 15:23 jnuche@deploy1003: jnuche, jforrester: Backport for [[gerrit:1342279{{!}}Use parser output value instead of status (T438154)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:19 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie
* 15:18 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir7003.magru.wmnet with reason: host reimage
* 15:14 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir7003.magru.wmnet with reason: host reimage
* 15:03 jnuche@deploy1003: Started scap sync-world: Backport for [[gerrit:1342279{{!}}Use parser output value instead of status (T438154)]]
* 14:50 moritzm: installing apache2 security updates
* 14:49 slyngshede@cumin1003: conftool action : set/pooled=yes; selector: name=cp5026.eqsin.wmnet
* 14:47 slyngshede@cumin1003: conftool action : set/weight=1; selector: name=cp5026.eqsin.wmnet
* 14:45 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5026.eqsin.wmnet with OS trixie
* 14:44 moritzm: installing python-filelock security updates
* 14:42 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir7003.magru.wmnet with OS trixie
* 14:35 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:35 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:34 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:33 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp6002.drmrs.wmnet
* 14:32 sukhe@puppetserver1001: conftool action : set/weight=100; selector: name=cp6002.drmrs.wmnet,service=ats-be
* 14:32 sukhe@puppetserver1001: conftool action : set/weight=1; selector: name=cp6002.drmrs.wmnet,service=cdn
* 14:29 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6002.drmrs.wmnet with OS trixie
* 14:24 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: sync
* 14:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:24 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:24 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: sync
* 14:14 jforrester@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.19,1.47.0-wmf.20,next --multiversion-image-basename docker-registry.discovery.wmnet/restricte
* 14:14 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:14 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:13 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1342008{{!}}abstractwiki: Add three new articles per community advice to show off the feature (T434227)]]
* 14:10 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: sync
* 14:10 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: sync
* 14:10 jforrester@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.19,1.47.0-wmf.20,next --multiversion-image-basename docker-registry.discovery.wmnet/restricte
* 14:10 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: sync
* 14:10 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: sync
* 14:09 Amir1: dropped links tables on db2237 ([[phab:T437278|T437278]])
* 14:08 Amir1: dropped links tables on db1238 ([[phab:T437278|T437278]])
* 14:07 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1342008{{!}}abstractwiki: Add three new articles per community advice to show off the feature (T434227)]]
* 14:03 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5026.eqsin.wmnet with reason: host reimage
* 14:02 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:02 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:02 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 14:01 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:00 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 13:59 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6002.drmrs.wmnet with reason: host reimage
* 13:56 slyngshede@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5026.eqsin.wmnet with reason: host reimage
* 13:54 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cp6002.drmrs.wmnet with reason: host reimage
* 13:53 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 13:52 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 13:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dse-k8s-etcd[1001-1003].eqiad.wmnet
* 13:48 bking@cumin2003: START - Cookbook sre.hosts.remove-downtime for dse-k8s-etcd[1001-1003].eqiad.wmnet
* 13:46 bking@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM dse-k8s-etcd1001.eqiad.wmnet
* 13:46 bking@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM dse-k8s-etcd1001.eqiad.wmnet
* 13:45 bking@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM dse-k8s-etcd1002.eqiad.wmnet
* 13:41 bking@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM dse-k8s-etcd1002.eqiad.wmnet
* 13:41 bking@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM dse-k8s-etcd1003.eqiad.wmnet
* 13:38 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp6002.drmrs.wmnet with OS trixie
* 13:37 bking@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM dse-k8s-etcd1003.eqiad.wmnet
* 13:37 bking@cumin2003: END (FAIL) - Cookbook sre.ganeti.reboot-vm (exit_code=99) for VM dse-k8s-etcd1003.eqiad.wmnet
* 13:37 bking@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM dse-k8s-etcd1003.eqiad.wmnet
* 13:34 slyngshede@cumin1003: START - Cookbook sre.hosts.reimage for host cp5026.eqsin.wmnet with OS trixie
* 13:34 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 13:33 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5026.mgmt.eqsin.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:29 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp6002.mgmt.drmrs.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:25 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on dse-k8s-etcd[1001-1003].eqiad.wmnet with reason: Maintenance to increase vCPUS [[phab:T438084|T438084]]
* 13:24 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:24 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:22 slyngshede@cumin1003: START - Cookbook sre.hosts.provision for host cp5026.mgmt.eqsin.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342231{{!}}SI: Unset all filters on links to cases (T434530)]], [[gerrit:1342234{{!}}SI: Unset all filters on links to cases (T434530)]] (duration: 13m 10s)
* 13:19 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org
* 13:19 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org
* 13:19 sukhe@cumin1004: START - Cookbook sre.hosts.provision for host cp6002.mgmt.drmrs.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:18 stran@deploy1003: stran: Continuing with deployment
* 13:15 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: apply
* 13:15 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: apply
* 13:13 stran@deploy1003: stran: Backport for [[gerrit:1342231{{!}}SI: Unset all filters on links to cases (T434530)]], [[gerrit:1342234{{!}}SI: Unset all filters on links to cases (T434530)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:11 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp5026.eqsin.wmnet
* 13:09 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1342231{{!}}SI: Unset all filters on links to cases (T434530)]], [[gerrit:1342234{{!}}SI: Unset all filters on links to cases (T434530)]]
* 13:09 sukhe@cumin1004: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp6002.drmrs.wmnet
* 13:05 sukhe@cumin1004: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp6002.drmrs.wmnet
* 13:05 sukhe@cumin1004: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts cp6002.drmrs.wmnet
* 13:00 dkertesz@cumin1004: conftool action : set/pooled=yes; selector: name=cp5025.eqsin.wmnet
* 12:59 dkertesz@cumin1004: conftool action : set/weight=1; selector: name=cp5025.eqsin.wmnet
* 12:52 sukhe@cumin1004: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp6002.drmrs.wmnet
* 12:52 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cp6002.mgmt.drmrs.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:51 dkertesz: eqsin pooled again ([[phab:T438052|T438052]])
* 12:49 dkertesz@cumin1004: conftool action : set/pooled=yes; selector: cluster=dnsbox,dc=eqsin
* 12:47 dkertesz@dns1004: END - running authdns-update
* 12:45 dkertesz@dns1004: START - running authdns-update
* 12:43 dkertesz@cumin1004: conftool action : set/pooled=yes; selector: cluster=dnsbox,dc=eqsin,service=authdns-update
* 12:41 dkertesz@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqsin [reason: no reason specified, [[phab:T438052|T438052]]]
* 12:41 dkertesz@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool eqsin [reason: no reason specified, [[phab:T438052|T438052]]]
* 12:38 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a1-eqiad
* 12:38 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a1-eqiad
* 12:34 sukhe@cumin1004: START - Cookbook sre.hosts.provision for host cp6002.mgmt.drmrs.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:34 sukhe@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cp6002.drmrs.wmnet with reason: reimage
* 12:33 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp6002.drmrs.wmnet
* 12:13 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5025.eqsin.wmnet with OS trixie
* 12:12 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 12:11 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-a1-eqiad
* 12:09 cmooney@cumin1004: START - Cookbook sre.network.tls for network device ssw1-a1-eqiad
* 12:01 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.20 refs [[phab:T430839|T430839]]
* 11:59 moritzm: pruned obsolete Bullseye image python3-bullseye from the docker registry [[phab:T416452|T416452]]
* 11:50 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341284{{!}}IS/IS-labs: Set wmgUseModeratorToolkit default false (T431000)]] (duration: 10m 32s)
* 11:46 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ml-lab1002.eqiad.wmnet
* 11:45 samtar@deploy1003: samtar: Continuing with deployment
* 11:44 samtar@deploy1003: samtar: Backport for [[gerrit:1341284{{!}}IS/IS-labs: Set wmgUseModeratorToolkit default false (T431000)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:41 klausman@cumin1003: START - Cookbook sre.hosts.reboot-single for host ml-lab1002.eqiad.wmnet
* 11:39 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1341284{{!}}IS/IS-labs: Set wmgUseModeratorToolkit default false (T431000)]]
* 11:39 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5025.eqsin.wmnet with reason: host reimage
* 11:35 slyngshede@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5025.eqsin.wmnet with reason: host reimage
* 11:34 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:33 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:31 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:31 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:27 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:27 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:26 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:25 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:24 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.20 refs [[phab:T430839|T430839]]
* 11:21 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply
* 11:21 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply
* 11:20 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply
* 11:20 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply
* 11:19 moritzm: kicked off a new run of production-images-weekly-rebuild.service on build2004 (previously some leftovers of buster in the config prevented a complete run)
* 11:17 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
* 11:16 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
* 11:11 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
* 11:11 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
* 11:10 jnuche@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342210{{!}}Revert "Tell VisualEditor about the app web edit tags" (T437736 T438125)]] (duration: 33m 14s)
* 11:10 slyngshede@cumin1003: START - Cookbook sre.hosts.reimage for host cp5025.eqsin.wmnet with OS trixie
* 11:05 marostegui@cumin1004: dbctl commit (dc=all): 'Remove db1180 from dbctl [[phab:T437222|T437222]]', diff saved to https://phabricator.wikimedia.org/P96459 and previous config saved to /var/cache/conftool/dbconfig/20260916-110502-marostegui.json
* 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 11:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:01 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 11:01 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 10:57 jnuche@deploy1003: jnuche: Continuing with deployment
* 10:57 jnuche@deploy1003: jnuche: Backport for [[gerrit:1342210{{!}}Revert "Tell VisualEditor about the app web edit tags" (T437736 T438125)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:37 jnuche@deploy1003: Started scap sync-world: Backport for [[gerrit:1342210{{!}}Revert "Tell VisualEditor about the app web edit tags" (T437736 T438125)]]
* 10:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5025.mgmt.eqsin.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:10 slyngshede@cumin1003: START - Cookbook sre.hosts.provision for host cp5025.mgmt.eqsin.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:02 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 10:02 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 09:58 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 09:49 slyngshede@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on cp5025.eqsin.wmnet with reason: reimaging
* 09:48 slyngshede@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cp5025.eqsin.wmnet with reason: reimaging
* 09:41 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-timeline: apply
* 09:41 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-timeline: apply
* 09:38 moritzm: imported routinator 0.15.2-1trixie to thirdparty/routinator [[phab:T438122|T438122]]
* 09:30 slyngshede@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cp5025.mgmt.eqsin.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:29 slyngshede@cumin1003: START - Cookbook sre.hosts.provision for host cp5025.mgmt.eqsin.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:19 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 09:19 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp5025.eqsin.wmnet
* 09:19 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'sync'.
* 09:18 slyngshede@cumin1003: conftool action : set/pooled=yes; selector: name=cp3074.esams.wmnet
* 09:18 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp3074.esams.wmnet
* 09:15 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 09:12 elukey@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'sync'.
* 09:12 elukey@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'sync'.
* 09:11 elukey@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'sync'.
* 09:11 elukey@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'sync'.
* 09:10 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 09:10 elukey@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 09:09 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 09:09 elukey@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 08:55 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 08:54 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 08:40 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:40 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:36 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqsin [reason: depooling for maintainance, [[phab:T438052|T438052]]]
* 08:36 slyngshede@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool eqsin [reason: depooling for maintainance, [[phab:T438052|T438052]]]
* 08:35 slyngshede@cumin1003: END (FAIL) - Cookbook sre.dns.admin (exit_code=99) DNS admin: depool eqsin [reason: no reason specified, no task ID specified]
* 08:35 slyngshede@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool eqsin [reason: no reason specified, no task ID specified]
* 08:35 slyngshede@cumin1003: conftool action : set/pooled=no; selector: cluster=dnsbox,dc=eqsin
* 08:34 fabfur: start depooling eqsin ([[phab:T438052|T438052]])
* 08:24 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.20 refs [[phab:T430839|T430839]]
* 08:22 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:22 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:14 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.20 refs [[phab:T430839|T430839]]
* 08:11 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 08:11 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 08:11 jnuche@deploy1003: Rolling back deployment
* 08:10 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 08:07 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 07:59 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 07:59 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 07:58 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 07:54 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 07:34 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply
* 07:34 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply
* 07:20 mlitn@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342112{{!}}Adds an instrument for pre-image-carousel-retest (T437076)]], [[gerrit:1342113{{!}}Adds an instrument for pre-image-carousel-retest (T437076)]], [[gerrit:1342117{{!}}Set up instrument for 5-arm test (T437076)]], [[gerrit:1342118{{!}}Set up instrument for 5-arm test (T437076)]] (duration: 10m 56s)
* 07:16 mlitn@deploy1003: mlitn: Continuing with deployment
* 07:15 mlitn@deploy1003: mlitn: Backport for [[gerrit:1342112{{!}}Adds an instrument for pre-image-carousel-retest (T437076)]], [[gerrit:1342113{{!}}Adds an instrument for pre-image-carousel-retest (T437076)]], [[gerrit:1342117{{!}}Set up instrument for 5-arm test (T437076)]], [[gerrit:1342118{{!}}Set up instrument for 5-arm test (T437076)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be veri
* 07:09 mlitn@deploy1003: Started scap sync-world: Backport for [[gerrit:1342112{{!}}Adds an instrument for pre-image-carousel-retest (T437076)]], [[gerrit:1342113{{!}}Adds an instrument for pre-image-carousel-retest (T437076)]], [[gerrit:1342117{{!}}Set up instrument for 5-arm test (T437076)]], [[gerrit:1342118{{!}}Set up instrument for 5-arm test (T437076)]]
* 06:50 moritzm: installing sudo security updates
* 06:47 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply
* 06:37 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply
* 05:12 moritzm: pruned obsolete Bullseye image buildkitd from the docker registry [[phab:T416452|T416452]]
* 04:56 kart_: Updated Apertium to 2026-09-15-084320-production ([[phab:T437213|T437213]])
* 04:54 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/apertium: apply
* 04:54 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/apertium: apply
* 04:50 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/apertium: apply
* 04:49 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/apertium: apply
* 04:45 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/apertium: apply
* 04:45 kartik@deploy1003: helmfile [staging] START helmfile.d/services/apertium: apply
* 04:24 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 04:24 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 04:22 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 04:22 oblivian@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 36s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-15 ==
* 23:09 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply
* 23:08 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply
* 23:08 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply
* 23:08 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply
* 23:07 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 23:07 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 23:07 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 23:07 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 23:06 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 23:06 rzl@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 22:57 sukhe@puppetserver1001: conftool action : set/weight=1; selector: name=cp6001.drmrs.wmnet,service=cdn
* 22:50 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc-wf1002.eqiad.wmnet with OS trixie
* 22:46 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs ([[phab:T436363|T436363]])
* 22:46 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs6003.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:46 brett@cumin2003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs6003.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:46 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs6003.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:46 brett@cumin2003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs6003.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:45 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs6002.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:45 brett@cumin2003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs6002.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:45 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs6002.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:44 brett@cumin2003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs6002.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:44 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs6001.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:44 brett@cumin2003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs6001.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:43 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs6001.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:43 brett@cumin2003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs6001.drmrs.wmnet<nowiki>}</nowiki> and A:liberica
* 22:43 brett@cumin2003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs ([[phab:T436363|T436363]])
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc-wf1002.eqiad.wmnet with reason: host reimage
* 22:33 brett@puppetserver1001: conftool action : set/weight=100; selector: name=cp6001.*
* 22:32 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp6001.*
* 22:26 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc-wf1002.eqiad.wmnet with reason: host reimage
* 22:07 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc-wf1002
* 22:07 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-wf1002
* 22:07 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host mc-wf1002
* 22:07 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc-wf1002.eqiad.wmnet 142.48.64.10.in-addr.arpa 2.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:07 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache mc-wf1002.eqiad.wmnet 142.48.64.10.in-addr.arpa 2.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:07 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:07 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc-wf1002 - rzl@cumin2003"
* 22:07 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc-wf1002 - rzl@cumin2003"
* 22:02 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:01 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host mc-wf1002
* 22:01 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host mc-wf1002.eqiad.wmnet with OS trixie
* 21:57 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6001.drmrs.wmnet with OS trixie
* 21:55 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply
* 21:55 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply
* 21:53 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply
* 21:53 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply
* 21:53 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:53 rzl@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:52 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:52 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:48 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:48 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:34 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6001.drmrs.wmnet with reason: host reimage
* 21:30 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp6001.drmrs.wmnet with reason: host reimage
* 21:20 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342049{{!}}MobileFrontend: Add app icons (T434258)]] (duration: 11m 47s)
* 21:15 jdlrobson@deploy1003: jdlrobson, cklimas: Continuing with deployment
* 21:13 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp6001.drmrs.wmnet with OS trixie
* 21:12 jdlrobson@deploy1003: jdlrobson, cklimas: Backport for [[gerrit:1342049{{!}}MobileFrontend: Add app icons (T434258)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:12 brett@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp6001.mgmt.drmrs.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:08 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1342049{{!}}MobileFrontend: Add app icons (T434258)]]
* 20:51 cdobbins@puppetserver1001: conftool action : set/weight=1; selector: name=cp2046.codfw.wmnet
* 20:51 cdobbins@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp2046.codfw.wmnet
* 20:48 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341271{{!}}Parsoid Read Views: Enable on all namespaces on wikitech (labswiki) (T437916)]] (duration: 09m 11s)
* 20:47 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp2046.codfw.wmnet with OS trixie
* 20:43 arlolra@deploy1003: ssastry, arlolra: Continuing with deployment
* 20:42 arlolra@deploy1003: ssastry, arlolra: Backport for [[gerrit:1341271{{!}}Parsoid Read Views: Enable on all namespaces on wikitech (labswiki) (T437916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:38 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1341271{{!}}Parsoid Read Views: Enable on all namespaces on wikitech (labswiki) (T437916)]]
* 20:34 brett@cumin2003: START - Cookbook sre.hosts.provision for host cp6001.mgmt.drmrs.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:30 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp6001.*
* 20:24 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp2046.codfw.wmnet with reason: host reimage
* 20:23 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339741{{!}}Enable ReaderExperiments in eswiki, jawiki, and ptwiki (T438009)]] (duration: 15m 58s)
* 20:20 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cp2046.codfw.wmnet with reason: host reimage
* 20:19 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T436363|T436363]])
* 20:19 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs4010.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:19 brett@cumin2003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs4010.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:19 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs4010.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:19 brett@cumin2003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs4010.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:18 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs4009.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:18 arlolra@deploy1003: lwatson, arlolra: Continuing with deployment
* 20:18 brett@cumin2003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs4009.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:17 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs4009.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:17 brett@cumin2003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs4009.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:17 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs4008.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:17 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp3075.esams.wmnet
* 20:17 sukhe@puppetserver1001: conftool action : set/weight=1; selector: name=cp3075.esams.wmnet
* 20:17 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp1103.eqiad.wmnet
* 20:17 brett@cumin2003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs4008.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:16 sukhe@puppetserver1001: conftool action : set/weight=1; selector: name=cp1103.eqiad.wmnet
* 20:16 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs4008.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:16 brett@cumin2003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs4008.ulsfo.wmnet<nowiki>}</nowiki> and A:liberica
* 20:16 brett@cumin2003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T436363|T436363]])
* 20:11 arlolra@deploy1003: lwatson, arlolra: Backport for [[gerrit:1339741{{!}}Enable ReaderExperiments in eswiki, jawiki, and ptwiki (T438009)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:10 brett@cumin2003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) config_reloading A:liberica-ulsfo ([[phab:T436363|T436363]])
* 20:08 brett@cumin2003: START - Cookbook sre.loadbalancer.admin config_reloading A:liberica-ulsfo ([[phab:T436363|T436363]])
* 20:08 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp1103.eqiad.wmnet with OS trixie
* 20:07 inflatador: bking@ganeti1046 sudo gnt-instance modify -B memory=4g,vcpus=4 dse-k8s-etcd100[1-3].eqiad.wmnet [[phab:T438084|T438084]]
* 20:07 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1339741{{!}}Enable ReaderExperiments in eswiki, jawiki, and ptwiki (T438009)]]
* 20:06 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp3075.esams.wmnet with OS trixie
* 20:04 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp7009.*
* 20:04 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host cp2046.codfw.wmnet with OS trixie
* 20:02 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp7009.*
* 20:02 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp7009.*
* 20:02 brett@puppetserver1001: conftool action : set/weight=1; selector: name=cp7009.*
* 20:01 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp7009.magru.wmnet with OS trixie
* 19:53 brett@puppetserver1001: conftool action : set/weight=1; selector: name=cp4045.*
* 19:53 brett@puppetserver1001: conftool action : set/weight=1; selector: name=cp4046.*
* 19:52 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4046.*
* 19:51 cdobbins@puppetserver1001: conftool action : set/pooled=no; selector: name=cp2046.codfw.wmnet
* 19:51 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp4046.ulsfo.wmnet with OS trixie
* 19:49 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4045.*
* 19:45 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp1103.eqiad.wmnet with reason: host reimage
* 19:43 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp4045.ulsfo.wmnet with OS trixie
* 19:41 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp3075.esams.wmnet with reason: host reimage
* 19:39 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cp1103.eqiad.wmnet with reason: host reimage
* 19:37 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp7009.magru.wmnet with reason: host reimage
* 19:33 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cp3075.esams.wmnet with reason: host reimage
* 19:32 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp7009.magru.wmnet with reason: host reimage
* 19:27 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp4046.ulsfo.wmnet with reason: host reimage
* 19:23 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp4046.ulsfo.wmnet with reason: host reimage
* 19:21 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp1103.eqiad.wmnet with OS trixie
* 19:19 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp1103.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:19 sukhe@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cp1103.eqiad.wmnet with reason: reimage
* 19:18 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp4045.ulsfo.wmnet with reason: host reimage
* 19:13 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 19:12 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 19:12 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 19:12 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp3075.esams.wmnet with OS trixie
* 19:12 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp4045.ulsfo.wmnet with reason: host reimage
* 19:12 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 19:11 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp3075.mgmt.esams.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:10 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp7009.magru.wmnet with OS trixie
* 19:09 brett@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp7009.mgmt.magru.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:08 sukhe@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cp3075.esams.wmnet with reason: reimaging
* 19:08 sukhe@cumin1004: START - Cookbook sre.hosts.provision for host cp1103.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:05 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp4046.ulsfo.wmnet with OS trixie
* 19:05 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore1006.eqiad.wmnet with OS bookworm
* 19:04 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3075.*
* 19:02 brett@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp4046.mgmt.ulsfo.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:01 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp1103.eqiad.wmnet
* 19:01 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp1103.eqiad.wmnet
* 19:00 sukhe@cumin1004: START - Cookbook sre.hosts.provision for host cp3075.mgmt.esams.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:58 brett@cumin2003: START - Cookbook sre.hosts.provision for host cp7009.mgmt.magru.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:57 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3075.esams.wmnet
* 18:55 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp1101.eqiad.wmnet
* 18:55 sukhe@puppetserver1001: conftool action : set/weight=1; selector: name=cp1101.eqiad.wmnet
* 18:55 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp4045.ulsfo.wmnet with OS trixie
* 18:52 brett@cumin2003: START - Cookbook sre.hosts.provision for host cp4046.mgmt.ulsfo.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:52 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp1101.eqiad.wmnet with OS trixie
* 18:45 cdobbins@puppetserver1001: conftool action : set/weight=1; selector: name=cp2044.codfw.wmnet
* 18:44 cdobbins@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp2044.codfw.wmnet
* 18:44 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore1006.eqiad.wmnet with reason: host reimage
* 18:40 brett@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp4045.mgmt.ulsfo.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:40 eevans@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore1006.eqiad.wmnet with reason: host reimage
* 18:39 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp3074.esams.wmnet
* 18:36 sukhe@cumin1004: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) config_reloading P<nowiki>{</nowiki>lvs3008.esams.wmnet<nowiki>}</nowiki> and A:liberica
* 18:36 sukhe@cumin1004: START - Cookbook sre.loadbalancer.admin config_reloading P<nowiki>{</nowiki>lvs3008.esams.wmnet<nowiki>}</nowiki> and A:liberica
* 18:33 sukhe@puppetserver1001: conftool action : set/weight=1; selector: name=cp3074.esams.wmnet
* 18:32 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp2044.codfw.wmnet with OS trixie
* 18:31 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp3074.esams.wmnet with OS trixie
* 18:30 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp1101.eqiad.wmnet with reason: host reimage
* 18:29 brett@cumin2003: START - Cookbook sre.hosts.provision for host cp4045.mgmt.ulsfo.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:26 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cp1101.eqiad.wmnet with reason: host reimage
* 18:22 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp7009.magru.wmnet with OS trixie
* 18:20 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1006.eqiad.wmnet with OS bookworm
* 18:20 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:19 eevans@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:19 eevans@cumin1004: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore1006.eqiad.wmnet
* 18:19 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore1006.eqiad.wmnet
* 18:10 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp1101.eqiad.wmnet with OS trixie
* 18:09 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp2044.codfw.wmnet with reason: host reimage
* 18:09 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp4045.ulsfo.wmnet with OS trixie
* 18:08 eevans@cumin1004: START - Cookbook sre.hosts.reboot-single for host sessionstore1006.eqiad.wmnet
* 18:07 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp3074.esams.wmnet with reason: host reimage
* 17:52 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp7009.magru.wmnet with reason: host reimage
* 17:48 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp2044.codfw.wmnet with OS trixie
* 17:43 cdobbins@puppetserver1001: conftool action : set/pooled=no; selector: name=cp2044.codfw.wmnet
* 17:36 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp3074.esams.wmnet with OS trixie
* 17:34 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp4046.ulsfo.wmnet with OS trixie
* 17:34 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp4045.ulsfo.wmnet with OS trixie
* 17:33 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3074.esams.wmnet
* 17:28 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp4046.*
* 17:28 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp4045.*
* 17:26 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp7009.magru.wmnet with OS trixie
* 17:25 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp1101.eqiad.wmnet with OS trixie
* 17:24 cdobbins@puppetserver1001: conftool action : set/pooled=no; selector: name=cp7009.magru.wmnet
* 17:23 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp1101.eqiad.wmnet
* 17:22 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:22 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org
* 17:02 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp7009.*
* 17:01 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:01 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:00 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:59 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:59 vriley@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host ms-be1099
* 16:59 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ms-be1099
* 16:59 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:56 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 16:55 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:55 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:55 vriley@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host ms-be1099
* 16:55 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ms-be1099
* 16:55 vriley@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host ms-be1099
* 16:54 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ms-be1099
* 16:54 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1098.eqiad.wmnet with OS bullseye
* 16:53 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:53 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [ms-be1099] - vriley@cumin1003"
* 16:53 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [ms-be1099] - vriley@cumin1003"
* 16:49 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 16:33 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS bullseye
* 16:21 mutante: temp disabling puppet on C:zookeeper (32 hosts) - safe deploy of https://gerrit.wikimedia.org/r/c/operations/puppet/+/1327569
* 16:04 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341890{{!}}Restore table borders for client-side MathJax (T435274)]], [[gerrit:1340558{{!}}lift IP cap for edit-a-thon /workshop (T437609 T437594 T437470)]] (duration: 24m 19s)
* 15:59 krinkle@deploy1003: anzx, krinkle: Continuing with deployment
* 15:44 krinkle@deploy1003: anzx, krinkle: Backport for [[gerrit:1341890{{!}}Restore table borders for client-side MathJax (T435274)]], [[gerrit:1340558{{!}}lift IP cap for edit-a-thon /workshop (T437609 T437594 T437470)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:40 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1341890{{!}}Restore table borders for client-side MathJax (T435274)]], [[gerrit:1340558{{!}}lift IP cap for edit-a-thon /workshop (T437609 T437594 T437470)]]
* 15:34 elukey@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 15:34 elukey@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 15:33 brennen@deploy1003: Finished deploy [phabricator/deployment@c386249]: deploy phab1005 for [[phab:T437930|T437930]] (duration: 00m 39s)
* 15:33 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS bullseye
* 15:33 brennen@deploy1003: Started deploy [phabricator/deployment@c386249]: deploy phab1005 for [[phab:T437930|T437930]]
* 15:32 elukey@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 15:32 elukey@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'sync'.
* 15:32 brennen@deploy1003: Finished deploy [phabricator/deployment@c386249]: deploy phab2003 for [[phab:T437930|T437930]] (duration: 00m 52s)
* 15:32 elukey@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 15:32 elukey@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'sync'.
* 15:31 brennen@deploy1003: Started deploy [phabricator/deployment@c386249]: deploy phab2003 for [[phab:T437930|T437930]]
* 15:31 elukey@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'.
* 15:31 elukey@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'.
* 15:26 jelto@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on phab2003.codfw.wmnet,phab[1005-1006].eqiad.wmnet with reason: Phabricator deploy
* 15:26 moritzm: pruned obsolete Bullseye image amd-gpu-tester from the docker registry [[phab:T416452|T416452]]
* 15:12 elukey@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 15:12 elukey@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 15:11 elukey@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 15:11 elukey@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 15:00 tgr@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341292{{!}}CommonSettings: Use a restrictive CSP for auth.wikimedia.org (T419684)]] (duration: 25m 11s)
* 14:55 tgr@deploy1003: tgr, arendpieter: Continuing with deployment
* 14:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:53 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:53 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:52 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:49 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 14:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 14:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 14:47 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 14:47 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 14:45 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 14:45 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 14:45 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 14:44 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 14:42 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 14:42 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 14:39 tgr@deploy1003: tgr, arendpieter: Backport for [[gerrit:1341292{{!}}CommonSettings: Use a restrictive CSP for auth.wikimedia.org (T419684)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:34 tgr@deploy1003: Started scap sync-world: Backport for [[gerrit:1341292{{!}}CommonSettings: Use a restrictive CSP for auth.wikimedia.org (T419684)]]
* 14:17 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341895{{!}}ReportIncidentController: Instance cache expensive methods (T437588)]] (duration: 11m 56s)
* 14:16 btullis@cumin1004: END (PASS) - Cookbook sre.ceph.rotate-osd-keys (exit_code=0) rolling rotate_keys on A:cephosd-codfw
* 14:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:12 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 14:09 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1341895{{!}}ReportIncidentController: Instance cache expensive methods (T437588)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:05 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1341895{{!}}ReportIncidentController: Instance cache expensive methods (T437588)]]
* 13:43 btullis@cumin1004: START - Cookbook sre.ceph.rotate-osd-keys rolling rotate_keys on A:cephosd-codfw
* 13:36 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341861{{!}}SuggestedInvestigations: Update "sockpuppet" queue view defaults (T438018)]] (duration: 10m 23s)
* 13:32 stran@deploy1003: stran: Continuing with deployment
* 13:30 stran@deploy1003: stran: Backport for [[gerrit:1341861{{!}}SuggestedInvestigations: Update "sockpuppet" queue view defaults (T438018)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:26 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1341861{{!}}SuggestedInvestigations: Update "sockpuppet" queue view defaults (T438018)]]
* 13:21 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:20 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:19 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334946{{!}}ArticleGuidance: Remove the experiment configuration keys (T434487)]] (duration: 09m 19s)
* 13:16 btullis@cumin1004: END (PASS) - Cookbook sre.ceph.rotate-osd-keys (exit_code=0) rolling rotate_keys on P<nowiki>{</nowiki>cephosd2001.codfw.wmnet<nowiki>}</nowiki> and (A:cephosd-codfw or A:cephosd-eqiad)
* 13:15 sbisson@deploy1003: sbisson: Continuing with deployment
* 13:14 sbisson@deploy1003: sbisson: Backport for [[gerrit:1334946{{!}}ArticleGuidance: Remove the experiment configuration keys (T434487)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) config_reloading P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:10 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1334946{{!}}ArticleGuidance: Remove the experiment configuration keys (T434487)]]
* 13:10 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin config_reloading P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:08 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:08 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:07 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:07 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:07 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:07 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 13:03 btullis@cumin1004: START - Cookbook sre.ceph.rotate-osd-keys rolling rotate_keys on P<nowiki>{</nowiki>cephosd2001.codfw.wmnet<nowiki>}</nowiki> and (A:cephosd-codfw or A:cephosd-eqiad)
* 13:00 btullis@cumin1004: END (PASS) - Cookbook sre.ceph.rotate-osd-keys (exit_code=0) rolling rotate_keys on P<nowiki>{</nowiki>cephosd2001.codfw.wmnet<nowiki>}</nowiki> and (A:cephosd-codfw or A:cephosd-eqiad)
* 12:59 btullis@cumin1004: START - Cookbook sre.ceph.rotate-osd-keys rolling rotate_keys on P<nowiki>{</nowiki>cephosd2001.codfw.wmnet<nowiki>}</nowiki> and (A:cephosd-codfw or A:cephosd-eqiad)
* 12:46 btullis@cumin1004: END (PASS) - Cookbook sre.ceph.rotate-osd-keys (exit_code=0) rolling rotate_keys on P<nowiki>{</nowiki>cephosd2001.codfw.wmnet<nowiki>}</nowiki> and (A:cephosd-codfw or A:cephosd-eqiad)
* 12:45 btullis@cumin1004: START - Cookbook sre.ceph.rotate-osd-keys rolling rotate_keys on P<nowiki>{</nowiki>cephosd2001.codfw.wmnet<nowiki>}</nowiki> and (A:cephosd-codfw or A:cephosd-eqiad)
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: network maintenance finished, [[phab:T437984|T437984]]]
* 12:40 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: network maintenance finished, [[phab:T437984|T437984]]]
* 12:29 XioNoX: asw1-b4-magru> request system reboot - [[phab:T437984|T437984]]
* 12:24 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 12:24 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 12:24 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 12:23 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 12:23 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 12:23 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 12:22 moritzm: installing shadow security updates
* 12:19 slyngshede@puppetserver1001: conftool action : set/weight=1; selector: name=cp7010.magru.wmnet
* 12:13 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 12:13 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 12 hosts with reason: Switch maintenance
* 12:12 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-b4-magru,asw1-b4-magru IPv6,asw1-b4-magru.mgmt with reason: Switch maintenance
* 12:11 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 12:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: switch reboot, [[phab:T437984|T437984]]]
* 12:11 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: switch reboot, [[phab:T437984|T437984]]]
* 12:09 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on install7002.wikimedia.org with reason: switch reboot
* 12:08 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 12:07 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 12:07 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 12:07 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 12:06 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 12:06 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 12:03 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 12:03 XioNoX: push pfw policies - [[phab:T437627|T437627]]
* 12:01 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 12:01 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) config_reloading P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 12:00 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin config_reloading P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 11:56 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 11:56 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 11:33 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2250.codfw.wmnet with reason: cloning db2201
* 11:18 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7004.magru.wmnet
* 11:17 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7004.magru.wmnet
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 11:05 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 11:05 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 11:04 slyngshede@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart P<nowiki>{</nowiki>lvs7001.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 10:51 slyngshede@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp7010.magru.wmnet
* 10:34 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 10:33 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 10:30 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1098.eqiad.wmnet with OS trixie
* 10:21 moritzm: failover Ganeti master in magru to ganeti7001
* 10:20 moritzm: increased DRBD replication speed in Ganeti/magru [[phab:T428878|T428878]]
* 10:10 hashar@deploy1003: Finished deploy [integration/docroot@5cf09c8]: build: Updating npm dependencies (duration: 00m 13s)
* 10:10 hashar@deploy1003: Started deploy [integration/docroot@5cf09c8]: build: Updating npm dependencies
* 10:09 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 10:08 moritzm: increased DRBD replication speed in Ganeti/esams [[phab:T428878|T428878]]
* 10:07 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 10:05 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 10:05 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:39 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: switches reboot, [[phab:T437984|T437984]]]
* 09:39 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool esams [reason: switches reboot, [[phab:T437984|T437984]]]
* 09:31 XioNoX: asw1-by27-esams> request system reboot - [[phab:T437984|T437984]]
* 09:30 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS trixie
* 09:28 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp7010.magru.wmnet with OS trixie
* 09:26 ayounsi@cumin1003: END (FAIL) - Cookbook sre.network.depool-rack (exit_code=99) with action 'depool' for esams rack BY27
* 09:24 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for esams rack BY27
* 09:24 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1098.eqiad.wmnet with OS trixie
* 09:23 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS trixie
* 09:22 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be1098.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 09:15 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.20 refs [[phab:T430839|T430839]]
* 09:10 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1098.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 09:06 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-by27-esams,asw1-by27-esams IPv6,asw1-by27-esams.mgmt with reason: Switch maintenance
* 09:05 ayounsi@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on asw1-by27-esams IPv6,asw1-by27-esams.mgmt,asw1-by-27-esams with reason: Switch maintenance
* 09:04 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 12 hosts with reason: Switch maintenance
* 09:04 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp7010.magru.wmnet with reason: host reimage
* 09:01 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: switches reboot, [[phab:T437984|T437984]]]
* 09:00 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: switches reboot, [[phab:T437984|T437984]]]
* 09:00 slyngshede@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp7010.magru.wmnet with reason: host reimage
* 08:59 jnuche@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341698{{!}}RestSandbox: Pass JsonLocalizer instead of ResponseFactory to ModuleManager (T437982)]] (duration: 12m 03s)
* 08:55 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2197.codfw.wmnet with reason: cloning db2201
* 08:55 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on install3004.wikimedia.org with reason: switch reboot
* 08:53 jnuche@deploy1003: jnuche: Continuing with deployment
* 08:52 jnuche@deploy1003: jnuche: Backport for [[gerrit:1341698{{!}}RestSandbox: Pass JsonLocalizer instead of ResponseFactory to ModuleManager (T437982)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:49 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: sync
* 08:49 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: sync
* 08:47 jnuche@deploy1003: Started scap sync-world: Backport for [[gerrit:1341698{{!}}RestSandbox: Pass JsonLocalizer instead of ResponseFactory to ModuleManager (T437982)]]
* 08:33 slyngshede@cumin1003: START - Cookbook sre.hosts.reimage for host cp7010.magru.wmnet with OS trixie
* 08:26 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be1098.eqiad.wmnet with OS trixie
* 08:26 slyngshede@puppetserver1001: conftool action : set/pooled=no; selector: name=cp7010.magru.wmnet
* 08:25 XioNoX: asw1-b3-magru - Disable logging and file logging for BRCM_PKT - [[phab:T437984|T437984]]
* 08:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS trixie
* 08:18 dpogorzelski@dns1004: END - running authdns-update
* 08:15 dpogorzelski@dns1004: START - running authdns-update
* 08:14 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3005.esams.wmnet
* 08:13 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3005.esams.wmnet
* 08:07 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341274{{!}}SI: Implement "queue view" functionality (T437183)]], [[gerrit:1341242{{!}}SuggestedInvestigations: Add and enable 'sockpuppets' queue view (T437183)]], [[gerrit:1341254{{!}}Add wmf-specific Special:SuggestedInvestigations messages (T437183)]] (duration: 55m 27s)
* 07:54 stran@deploy1003: stran: Continuing with deployment
* 07:31 stran@deploy1003: stran: Backport for [[gerrit:1341274{{!}}SI: Implement "queue view" functionality (T437183)]], [[gerrit:1341242{{!}}SuggestedInvestigations: Add and enable 'sockpuppets' queue view (T437183)]], [[gerrit:1341254{{!}}Add wmf-specific Special:SuggestedInvestigations messages (T437183)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:18 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply
* 07:18 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply
* 07:11 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1341274{{!}}SI: Implement "queue view" functionality (T437183)]], [[gerrit:1341242{{!}}SuggestedInvestigations: Add and enable 'sockpuppets' queue view (T437183)]], [[gerrit:1341254{{!}}Add wmf-specific Special:SuggestedInvestigations messages (T437183)]]
* 07:06 moritzm: pruned obsolete Bullseye image python3-devel from the docker registry [[phab:T416452|T416452]]
* 06:51 moritzm: pruned obsolete Bullseye image python3-build-bullseye from the docker registry [[phab:T416452|T416452]]
* 05:59 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply
* 05:49 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply
* 05:48 oblivian@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 05:47 oblivian@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 05:38 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply
* 05:28 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-video: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-video: apply
* 05:10 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-video: apply
* 05:10 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-video: apply
* 05:10 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-video: apply
* 05:10 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-video: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-syntaxhighlight: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-syntaxhighlight: apply
* 05:10 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-syntaxhighlight: apply
* 05:10 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-syntaxhighlight: apply
* 05:10 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-syntaxhighlight: apply
* 05:10 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-syntaxhighlight: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-media: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-media: apply
* 05:10 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-media: apply
* 05:10 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-media: apply
* 05:10 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-media: apply
* 05:10 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-media: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-constraints: apply
* 05:10 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-constraints: apply
* 05:10 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-constraints: apply
* 05:09 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-constraints: apply
* 05:09 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-constraints: apply
* 05:09 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-constraints: apply
* 05:08 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply
* 05:08 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox: apply
* 05:07 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply
* 05:07 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox: apply
* 05:07 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox: apply
* 05:07 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox: apply
* 05:07 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply
* 05:07 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply
* 05:06 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox-timeline: apply
* 05:06 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox-timeline: apply
* 05:06 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply
* 05:06 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply
* 04:07 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.17 (duration: 07m 10s)
* 03:06 mwpresync@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.19,1.47.0-wmf.20,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.20 refs [[phab:T430839|T430839]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 22s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:43 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 00:23 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore1005.eqiad.wmnet with reason: host reimage
* 00:19 jclark@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore1005.eqiad.wmnet with reason: host reimage
* 00:17 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 00:07 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sessionstore1005.eqiad.wmnet with OS bookworm
== 2026-09-14 ==
* 23:41 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 23:26 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324966{{!}}Enable Produnto on pilot wikis (T421436)]] (duration: 12m 59s)
* 23:22 tstarling@deploy1003: tstarling: Continuing with deployment
* 23:17 tstarling@deploy1003: tstarling: Backport for [[gerrit:1324966{{!}}Enable Produnto on pilot wikis (T421436)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:13 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1324966{{!}}Enable Produnto on pilot wikis (T421436)]]
* 23:01 eevans@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 22:41 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341385{{!}}VisualEditor: don't register settings tool in wikitextCommandRegistry (T437810)]] (duration: 09m 22s)
* 22:36 kemayo@deploy1003: kemayo: Continuing with deployment
* 22:36 kemayo@deploy1003: kemayo: Backport for [[gerrit:1341385{{!}}VisualEditor: don't register settings tool in wikitextCommandRegistry (T437810)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:31 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1341385{{!}}VisualEditor: don't register settings tool in wikitextCommandRegistry (T437810)]]
* 22:22 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 22:22 eevans@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 21:43 sbassett: Deployed security fix for [[phab:T435623|T435623]]
* 21:29 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1098.eqiad.wmnet with OS bullseye
* 21:29 sbassett: Deployed security fix for [[phab:T434372|T434372]]
* 21:26 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 21:26 eevans@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 21:05 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338999{{!}}Enable ReadingLists for all logged-in users on English Wikipedia (T434923)]], [[gerrit:1340004{{!}}Enable Reading Recommendations experiment on test wiki (T437665)]] (duration: 11m 03s)
* 21:00 aude@deploy1003: aude, jdlrobson: Continuing with deployment
* 20:58 aude@deploy1003: aude, jdlrobson: Backport for [[gerrit:1338999{{!}}Enable ReadingLists for all logged-in users on English Wikipedia (T434923)]], [[gerrit:1340004{{!}}Enable Reading Recommendations experiment on test wiki (T437665)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:54 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1338999{{!}}Enable ReadingLists for all logged-in users on English Wikipedia (T434923)]], [[gerrit:1340004{{!}}Enable Reading Recommendations experiment on test wiki (T437665)]]
* 20:47 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341278{{!}}[A11y] Add list semantics to ReadingList page (T435864 T434923)]] (duration: 12m 49s)
* 20:43 aude@deploy1003: aude, jdlrobson: Continuing with deployment
* 20:39 aude@deploy1003: aude, jdlrobson: Backport for [[gerrit:1341278{{!}}[A11y] Add list semantics to ReadingList page (T435864 T434923)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:34 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1341278{{!}}[A11y] Add list semantics to ReadingList page (T435864 T434923)]]
* 20:32 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339811{{!}}wmf-config: Register content/v2-beta REST module as disabled (T432798)]], [[gerrit:1338274{{!}}wikifunctions: Move abstract fragments to mainstash (T432849)]] (duration: 25m 21s)
* 20:28 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:27 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:27 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:27 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:25 jforrester@deploy1003: jforrester, aghirelli: Continuing with deployment
* 20:24 jforrester@deploy1003: jforrester, aghirelli: Backport for [[gerrit:1339811{{!}}wmf-config: Register content/v2-beta REST module as disabled (T432798)]], [[gerrit:1338274{{!}}wikifunctions: Move abstract fragments to mainstash (T432849)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS bullseye
* 20:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be1098.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 20:06 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1339811{{!}}wmf-config: Register content/v2-beta REST module as disabled (T432798)]], [[gerrit:1338274{{!}}wikifunctions: Move abstract fragments to mainstash (T432849)]]
* 20:04 vriley@cumin1003: START - Cookbook sre.hosts.provision for host ms-be1098.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 20:03 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be1098
* 20:02 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ms-be1098
* 20:02 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 20:02 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [ms-be1098] - vriley@cumin1003"
* 20:02 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [ms-be1098] - vriley@cumin1003"
* 19:59 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 19:58 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 19:57 dzahn@dns1005: END - running authdns-update
* 19:55 eevans@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 19:55 dzahn@dns1005: START - running authdns-update
* 19:54 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 19:38 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341273{{!}}[CodeMirror] enable for new users (enwiki), new VE integration (global) (T288161 T432558)]] (duration: 33m 51s)
* 19:26 musikanimal@deploy1003: musikanimal: Continuing with deployment
* 19:22 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1341273{{!}}[CodeMirror] enable for new users (enwiki), new VE integration (global) (T288161 T432558)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:04 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1341273{{!}}[CodeMirror] enable for new users (enwiki), new VE integration (global) (T288161 T432558)]]
* 18:53 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:50 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 18:50 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 18:49 eevans@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 18:26 brett@cumin2003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d6-eqiad
* 18:26 brett@cumin2003: START - Cookbook sre.network.tls for network device lsw1-d6-eqiad
* 18:26 brett@cumin2003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-eqiad
* 18:26 brett@cumin2003: START - Cookbook sre.network.tls for network device ssw1-d8-eqiad
* 18:25 brett@cumin2003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-eqiad
* 18:25 brett@cumin2003: START - Cookbook sre.network.tls for network device lsw1-d4-eqiad
* 18:25 root@cumin2003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-eqiad
* 18:25 root@cumin2003: START - Cookbook sre.network.tls for network device lsw1-d2-eqiad
* 18:17 jclark@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 18:14 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337611{{!}}extension-list: Add ModeratorToolkit (T431000)]] (duration: 09m 34s)
* 18:10 samtar@deploy1003: samtar: Continuing with deployment
* 18:09 samtar@deploy1003: samtar: Backport for [[gerrit:1337611{{!}}extension-list: Add ModeratorToolkit (T431000)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:06 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 18:05 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1337611{{!}}extension-list: Add ModeratorToolkit (T431000)]]
* 18:01 jclark@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 17:48 jclark@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 17:07 tgr@deploy1003: Finished scap sync-world: Backport for [[gerrit:1330446{{!}}CommonSettings: Use a restrictive, eval-free CSP for auth.wikimedia.org (T419684)]] (duration: 23m 19s)
* 17:00 tgr@deploy1003: arendpieter, tgr: Rolling back deployment
* 16:52 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 16:51 eevans@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 16:49 tgr@deploy1003: arendpieter, tgr: Backport for [[gerrit:1330446{{!}}CommonSettings: Use a restrictive, eval-free CSP for auth.wikimedia.org (T419684)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:44 tgr@deploy1003: Started scap sync-world: Backport for [[gerrit:1330446{{!}}CommonSettings: Use a restrictive, eval-free CSP for auth.wikimedia.org (T419684)]]
* 16:09 Amir1: drop links tables from db1252 ([[phab:T437278|T437278]])
* 16:07 Amir1: drop links tables from db2240 ([[phab:T437278|T437278]])
* 16:05 Amir1: drop non-links tables from db2247 ([[phab:T437278|T437278]])
* 15:53 Lucas_WMDE: UTC afternoon backport+config window belatedly done
* 15:50 lucaswerkmeister-wmde@deploy1003: mwscript-k8s job started: namespaceDupes abstractwiki --fix # [[phab:T437772|T437772]]
* 15:49 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335401{{!}}Adjust extendedconfirmed calculation to first edit on viwiki (T437006)]], [[gerrit:1340505{{!}}core-Namespaces: Add AW and AWT alias for its talk in abstractwiki (T437772)]] (duration: 10m 23s)
* 15:48 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 15:47 eevans@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 15:45 lucaswerkmeister-wmde@deploy1003: bunnypranav, lucaswerkmeister-wmde, tryvix1509: Continuing with deployment
* 15:43 lucaswerkmeister-wmde@deploy1003: bunnypranav, lucaswerkmeister-wmde, tryvix1509: Backport for [[gerrit:1335401{{!}}Adjust extendedconfirmed calculation to first edit on viwiki (T437006)]], [[gerrit:1340505{{!}}core-Namespaces: Add AW and AWT alias for its talk in abstractwiki (T437772)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:39 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1335401{{!}}Adjust extendedconfirmed calculation to first edit on viwiki (T437006)]], [[gerrit:1340505{{!}}core-Namespaces: Add AW and AWT alias for its talk in abstractwiki (T437772)]]
* 15:36 elukey@dns1004: END - running authdns-update
* 15:36 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 15:35 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 15:35 moritzm: installing shadow security updates
* 15:35 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 15:34 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:33 elukey@dns1004: START - running authdns-update
* 15:33 eevans@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:32 eevans@cumin1004: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore1005.eqiad.wmnet
* 15:29 lucaswerkmeister-wmde@deploy1003: mwscript-k8s job started: namespaceDupes afwiki --fix # [[phab:T437576|T437576]]
* 15:29 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338902{{!}}afwiki: Create Draft and Draft talk namespaces (T437576)]] (duration: 15m 44s)
* 15:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 15:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 15:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 15:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 15:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 15:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 15:21 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, tryvix1509: Continuing with deployment
* 15:21 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:17 eevans@cumin1004: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore1005.eqiad.wmnet
* 15:17 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, tryvix1509: Backport for [[gerrit:1338902{{!}}afwiki: Create Draft and Draft talk namespaces (T437576)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:17 eevans@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore1005.eqiad.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 15:13 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore1004.eqiad.wmnet
* 15:13 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1338902{{!}}afwiki: Create Draft and Draft talk namespaces (T437576)]]
* 15:06 eevans@cumin1004: START - Cookbook sre.hosts.reboot-single for host sessionstore1004.eqiad.wmnet
* 14:59 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore1004.eqiad.wmnet with OS bookworm
* 14:38 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore1004.eqiad.wmnet with reason: host reimage
* 14:33 marostegui@dns1004: END - running authdns-update
* 14:32 eevans@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore1004.eqiad.wmnet with reason: host reimage
* 14:30 marostegui@dns1004: START - running authdns-update
* 14:15 eevans@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1004.eqiad.wmnet with OS bookworm
* 14:14 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore1004.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:13 eevans@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1004.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:13 eevans@cumin1004: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore1004.eqiad.wmnet
* 14:13 eevans@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore1004.eqiad.wmnet
* 14:05 moritzm: kick off a rebuild of base images on build2004
* 14:05 moritzm: kick off a rebuild of base images on build2004
* 14:00 eevans@cumin1004: START - Cookbook sre.hosts.reboot-single for host sessionstore1004.eqiad.wmnet
* 14:00 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:00 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 13:56 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:54 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:52 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:44 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 13:43 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 13:42 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:40 eevans@cumin1004: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore1004.eqiad.wmnet
* 13:40 eevans@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore1004.eqiad.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 13:40 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:40 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:39 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:39 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:39 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 13:38 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 13:38 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:38 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:38 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 13:37 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 13:37 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:36 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:36 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 13:35 oblivian@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:35 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 13:35 oblivian@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:35 oblivian@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:34 oblivian@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'.
* 13:34 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 13:34 oblivian@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:34 oblivian@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:34 oblivian@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:34 oblivian@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 13:34 oblivian@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'.
* 13:33 oblivian@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'.
* 13:33 oblivian@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
* 13:33 oblivian@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
* 13:33 oblivian@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 13:33 oblivian@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 13:33 oblivian@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 13:32 oblivian@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 13:32 oblivian@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 13:32 oblivian@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 13:32 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 13:32 oblivian@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 13:32 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 13:32 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 13:29 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 13:23 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1003.eqiad.wmnet
* 13:21 sukhe: sudo cumin -b11 "A:cp-text" "run-puppet-agent --enable 'merging CR 1338134'" [[phab:T425441|T425441]]
* 13:20 sukhe: sudo cumin -b11 "A:cp-text" "run-puppet-agent --enable 'merging CR 1338134'"[[phab:T425441|T425441]]
* 13:19 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 13:17 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1003.eqiad.wmnet
* 13:14 moritzm: installing Bird security updates
* 13:09 sukhe: sudo cumin "A:cp-text" "disable-puppet 'merging CR 1338134'"
* 13:06 jmm@dns1004: END - running authdns-update
* 13:04 jmm@dns1004: START - running authdns-update
* 12:58 moritzm: update Trixie installer image to 13.7 [[phab:T437715|T437715]]
* 12:58 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 12:54 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 12:52 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 12:49 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 12:48 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 12:47 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 12:44 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 12:44 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 12:42 oblivian@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 12:40 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply
* 12:40 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 12:40 oblivian@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 12:39 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 12:39 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 12:39 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 12:37 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 12:37 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 12:36 marostegui@cumin1004: conftool action : set/pooled=yes; selector: name=clouddb1025.eqiad.wmnet,service=x4
* 12:34 _joe_: adding gvisor labels to all wikikube clusters nodes
* 12:30 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply
* 12:14 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 12:14 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 11:22 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1260: After cloning
* 10:48 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:45 ladsgroup@dns1004: END - running authdns-update
* 10:42 ladsgroup@dns1004: START - running authdns-update
* 10:37 marostegui@cumin1004: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=x4
* 10:37 marostegui@cumin1004: START - Cookbook sre.mysql.pool pool db1260: After cloning
* 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 10:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 10:04 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 10:04 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 09:53 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:53 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:41 Amir1: drop links tables from db2172 ([[phab:T437278|T437278]])
* 09:40 Amir1: drop links tables from db1228 ([[phab:T437278|T437278]])
* 09:08 marostegui: Stop mariadb on db1260 to clone dbstore1007, there will be lag on wikireplicas:x4 https://phabricator.wikimedia.org/T437839
* 09:07 taavi@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1025.eqiad.wmnet
* 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1260: Needs to clone another host from this one
* 09:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1260: Needs to clone another host from this one
* 09:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb[1024-1025].eqiad.wmnet,db[1155,1260].eqiad.wmnet,dbstore1007.eqiad.wmnet with reason: Adding x4
* 08:44 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:42 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:41 moritzm: pruned obsolete Bullseye images php8.3-icu72-cli / php8.3-icu72-fpm-multiversion-base / php8.3-icu72-fpm from the docker registry [[phab:T416452|T416452]]
* 08:37 taavi@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet
* 08:37 taavi@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1024.eqiad.wmnet
* 08:36 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on dbstore1007.eqiad.wmnet with reason: Adding x4
* 08:35 moritzm: pruned obsolete Bullseye images php8.1-cli/php8.1-fpm/ php8.1-fpm-multiversion-base from the docker registry [[phab:T416452|T416452]]
* 08:10 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: sync
* 08:08 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: sync
* 07:58 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 23s)
* 07:57 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 07:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 07:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 07:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 07:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 07:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 07:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 07:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 07:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 07:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 07:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 07:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 07:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 07:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 07:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 07:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 07:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 07:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 07:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 07:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 07:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 07:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 07:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 07:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 07:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 07:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 07:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 07:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 07:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 07:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 07:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 07:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 07:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 07:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 07:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 07:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 07:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 07:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 07:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 07:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 07:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 07:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 07:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 07:31 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1340801{{!}}Allow title-like strings with Package: prefix in require() (T430644)]], [[gerrit:1340802{{!}}Runtime: Add a facility for loading files by title (T430644)]] (duration: 34m 30s)
* 07:18 tstarling@deploy1003: tstarling: Continuing with deployment
* 07:17 tstarling@deploy1003: tstarling: Backport for [[gerrit:1340801{{!}}Allow title-like strings with Package: prefix in require() (T430644)]], [[gerrit:1340802{{!}}Runtime: Add a facility for loading files by title (T430644)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 06:56 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1340801{{!}}Allow title-like strings with Package: prefix in require() (T430644)]], [[gerrit:1340802{{!}}Runtime: Add a facility for loading files by title (T430644)]]
* 06:26 TimStarling: on deploy1003: docker image pull docker-registry.wikimedia.org/php8.3-fpm-multiversion-base
* 05:51 _joe_: pulled bookworm:latest from build2004 to build2001 [[phab:T437829|T437829]]
* 05:39 _joe_: force-running build-base-images on build2004 for [[phab:T437829|T437829]]
* 04:53 TimStarling: on build2001 rebuilding base images [[phab:T437829|T437829]]
* 03:00 tstarling@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.18,1.47.0-wmf.19,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted
* 02:59 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1340801{{!}}Allow title-like strings with Package: prefix in require() (T430644)]], [[gerrit:1340802{{!}}Runtime: Add a facility for loading files by title (T430644)]]
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-13 ==
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 29s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-12 ==
* 19:40 ladsgroup@cumin1003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-eqiad
* 19:32 ladsgroup@cumin1003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-eqiad
* 19:30 ladsgroup@cumin1003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw
* 19:21 ladsgroup@cumin1003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 35s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-11 ==
* 21:51 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 21:50 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 16:47 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339808{{!}}Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813{{!}}Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]] (duration: 07m 23s)
* 16:43 sbassett@deploy1003: sbassett: Continuing with deployment
* 16:42 sbassett@deploy1003: sbassett: Backport for [[gerrit:1339808{{!}}Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813{{!}}Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:40 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1339808{{!}}Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813{{!}}Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]]
* 16:08 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 14:39 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 14:10 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:10 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:10 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:09 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 13:40 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:28 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:11 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:10 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: db1199 repool
* 11:05 moritzm: installing Linux 6.1.187 on Bookworm hosts
* 11:05 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Jcrespo out of all services on: 2443 hosts
* 10:44 aokoth@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2020 in turn
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2019 in turn
* 10:41 Emperor: restart versitygw@objectstorage0[0-3].service on backup2018 in turn
* 10:40 Emperor: restart versitygw@objectstorage0[0-3].service on backup2017 in turn
* 10:39 Emperor: restart versitygw@objectstorage0[0-3].service on backup2016 in turn
* 10:37 Emperor: restart versitygw@objectstorage0[0-3].service on backup2015 in turn
* 10:36 Emperor: restart versitygw@objectstorage0[0-3].service on backup1020 in turn
* 10:35 Emperor: restart versitygw@objectstorage0[0-3].service on backup1019 in turn
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1199: db1199 repool
* 10:33 Emperor: restart versitygw@objectstorage0[0-3].service on backup1018 in turn
* 10:32 Emperor: restart versitygw@objectstorage0[0-3].service on backup1017 in turn
* 10:30 Emperor: restart versitygw@objectstorage0[0-3].service on backup1016 in turn
* 10:20 Emperor: restart versitygw@objectstorage0[1-3].service on backup1015 in turn
* 10:17 Emperor: restart versitygw@objectstorage00.service on backup1015
* 10:15 aokoth@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 08:46 slyngs: Update CAS/SSO to CAS 7.3.8.3
* 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:45 slyngshede@dns1004: END - running authdns-update
* 08:43 slyngshede@dns1004: START - running authdns-update
* 08:36 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:28 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 7 hosts with reason: Restarting s5
* 08:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db[1154,1269].eqiad.wmnet with reason: Restarting s5
* 08:20 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:20 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Repooling db1159
* 07:58 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:58 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:54 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:28 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:27 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1199: Needs to clone another host from this one
* 07:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1199: Needs to clone another host from this one
* 07:16 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Repooling db1159
* 07:15 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1199.eqiad.wmnet with reason: Cloning s4
* 07:10 TimStarling: killed jobs for [[phab:T437056|T437056]] since they weren't purging
* 06:38 TimStarling: also started refreshLinks for ptwiki and zhwiki, reparsing ~3000 pages altogether [[phab:T437056|T437056]]
* 06:27 TimStarling: for [[phab:T437056|T437056]]: mwscript-k8s refreshLinks.php --wiki=eswiki --tracking-category scribunto-common-error-category
* 05:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1159: Needs to clone another host from this one
* 05:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1159: Needs to clone another host from this one
* 05:30 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1159.eqiad.wmnet with reason: Cloning
* 05:29 marostegui: Start cloning db1245:s5 [[phab:T437563|T437563]]
* 05:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet,db1245.eqiad.wmnet with reason: Cloning
* 05:25 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] (duration: 09m 59s)
* 05:21 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:20 tstarling@deploy1003: tstarling: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:15 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 50s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-10 ==
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1349.eqiad.wmnet
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1349.eqiad.wmnet
* 23:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1349.eqiad.wmnet
* 23:09 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1349
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1349
* 22:32 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1349
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:28 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1349
* 22:27 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:27 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1349.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1348.eqiad.wmnet
* 22:23 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 22:12 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] (duration: 10m 59s)
* 22:06 musikanimal@deploy1003: kemayo, musikanimal: Rolling back deployment
* 22:05 musikanimal@deploy1003: kemayo, musikanimal: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:01 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]]
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:00 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:52 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:47 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:47 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] (duration: 13m 23s)
* 21:46 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:46 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2006.codfw.wmnet
* 21:42 derenrich@deploy1003: derenrich: Continuing with deployment
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1348
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1348
* 21:39 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1348
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:39 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:37 derenrich@deploy1003: derenrich: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:35 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:34 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1348
* 21:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 21:33 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1348.eqiad.wmnet
* 21:33 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]]
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1348.eqiad.wmnet
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1348.eqiad.wmnet
* 21:31 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2006.codfw.wmnet
* 21:31 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] (duration: 09m 45s)
* 21:27 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:26 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]]
* 21:19 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:19 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2006.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 21:17 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] (duration: 13m 54s)
* 21:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2005.codfw.wmnet with OS bookworm
* 21:12 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:07 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdeb
* 21:03 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]]
* 20:54 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:52 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] (duration: 23m 49s)
* 20:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:47 jdrewniak@deploy1003: jdrewniak, milazg: Continuing with deployment
* 20:34 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1346.eqiad.wmnet
* 20:32 jdrewniak@deploy1003: jdrewniak, milazg: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2005.codfw.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:28 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]]
* 20:27 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2005.codfw.wmnet
* 20:26 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:24 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:22 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] (duration: 11m 24s)
* 20:17 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 20:15 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2005.codfw.wmnet
* 20:14 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:12 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 20:11 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:07 jdrewniak@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.18,1.47.0-wmf.19,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted
* 20:05 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:02 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:01 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2005.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 19:53 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:50 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1346
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 19:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1346.eqiad.wmnet
* 19:21 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 thcipriani: Gerrit downtime incoming for upgrade
* 19:17 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hadoop.roll-restart-workers (exit_code=0) restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 dzahn@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:30:00 on gerrit.wikimedia.org with reason: maintenance upgrade
* 19:16 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on gerrit2003.wikimedia.org with reason: maintenance upgrade
* 19:04 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 18:21 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e] (duration: 00m 59s)
* 18:20 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e]
* 18:19 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e] (duration: 05m 13s)
* 18:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 18:14 otto@deploy1003: Started deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e]
* 18:13 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e] (duration: 00m 39s)
* 18:13 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e]
* 18:13 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 18:12 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 18:11 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 18:11 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit2002.wikimedia.org with reason: maintenance upgrade
* 18:11 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:11 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 18:10 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit1003.wikimedia.org with reason: maintenance upgrade
* 18:09 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 18:08 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 18:06 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 16:40 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 16:35 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 16:33 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]] synced to the te
* 16:28 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 15:33 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4
* 15:04 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] (duration: 08m 08s)
* 15:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host clouddumps1001.wikimedia.org with OS bookworm
* 14:59 samtar@deploy1003: samtar: Continuing with deployment
* 14:58 samtar@deploy1003: samtar: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:56 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]]
* 14:40 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:07 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 13:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host clouddumps1001.wikimedia.org with OS bookworm
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:55 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:51 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:42 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:41 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:48 klausman@dns1004: END - running authdns-update
* 12:46 klausman@dns1004: START - running authdns-update
* 12:35 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 12:35 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 12:05 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning x4
* 12:01 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2005.codfw.wmnet
* 11:55 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2005.codfw.wmnet
* 11:54 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1144.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:52 cgoubert@dns1004: END - running authdns-update
* 11:49 cgoubert@dns1004: START - running authdns-update
* 11:31 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2004.codfw.wmnet
* 11:25 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2004.codfw.wmnet
* 11:24 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:16 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1200.eqiad.wmnet
* 11:16 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1200.eqiad.wmnet
* 11:04 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1200.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:04 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1199.eqiad.wmnet
* 11:03 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1199.eqiad.wmnet
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:42 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1199.eqiad.wmnet with reason: Upgrading RAID firmware
* 10:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1024.eqiad.wmnet with reason: Cloning x4
* 10:00 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 09:56 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 09:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1024.eqiad.wmnet
* 09:45 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 09:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 09:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:04 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:51 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:46 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning x4
* 08:43 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:34 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] (duration: 09m 56s)
* 08:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:30 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:29 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:28 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:24 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]]
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 08:09 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 07:58 XioNoX: netflow1004:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:56 XioNoX: netflow2005:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 07:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 07:03 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:59 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:43 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:42 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:39 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 06:38 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 06:37 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 05:04 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 05:03 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tool-server' for release 'main' .
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 38s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1345.eqiad.wmnet
* 00:11 dzahn@dns1004: END - running authdns-update
* 00:08 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 00:08 dzahn@dns1004: START - running authdns-update
== 2026-09-09 ==
* 23:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:29 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1345
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 23:28 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1345.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1344.eqiad.wmnet
* 23:22 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] (duration: 11m 15s)
* 23:18 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 23:16 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 23:11 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]]
* 22:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:51 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2004.codfw.wmnet with OS bookworm
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1344
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1344
* 22:37 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1344
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1344
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 22:32 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] (duration: 10m 21s)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1344.eqiad.wmnet
* 22:27 derenrich@deploy1003: derenrich, egardner: Continuing with deployment
* 22:26 derenrich@deploy1003: derenrich, egardner: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:24 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:22 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]]
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:19 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] (duration: 13m 40s)
* 22:16 derenrich@deploy1003: derenrich: Rolling back deployment
* 22:10 derenrich@deploy1003: derenrich: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:05 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]]
* 22:03 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:44 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:40 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:36 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1343
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1343
* 21:27 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1343
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:27 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:23 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:22 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1343
* 21:22 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] (duration: 12m 29s)
* 21:21 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 21:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1343.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1342.eqiad.wmnet
* 21:17 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:14 jforrester@deploy1003: jforrester: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:09 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]]
* 21:08 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:07 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 21:07 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:06 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2004.codfw.wmnet
* 20:59 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:57 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:56 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2004.codfw.wmnet
* 20:53 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:53 bking@cumin2003: END (ERROR) - Cookbook sre.presto.roll-restart-workers (exit_code=97) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:53 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:50 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* {{safesubst:SAL entry|1=20:45 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2}}
* 20:42 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* 20:41 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 20:40 sbassett@deploy1003: aranyap, sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:39 sbassett@deploy1003: aranyap, sbassett: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "Filter}}
* 20:35 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* {{safesubst:SAL entry|1=20:34 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "}}
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:28 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:24 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 20:24 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1342
* 20:23 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 20:23 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1342.eqiad.wmnet
* 20:23 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1342.eqiad.wmnet
* 20:22 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1342.eqiad.wmnet
* 20:15 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1027.eqiad.wmnet with OS bookworm
* 19:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:41 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1027.eqiad.wmnet with OS bookworm
* 19:36 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:28 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:26 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:19 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 19:19 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:06 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 19:05 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 19:03 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 18:59 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 18:23 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:06 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1026.eqiad.wmnet with OS bookworm
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 18:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 17:54 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 17:49 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 17:45 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1026.eqiad.wmnet with OS bookworm
* 17:31 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:29 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 17:27 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 17:12 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 17:04 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 16:46 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] (duration: 09m 28s)
* 16:41 urbanecm@deploy1003: migr, urbanecm: Continuing with deployment
* 16:41 urbanecm@deploy1003: migr, urbanecm: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:36 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]]
* 16:28 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:25 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 15:55 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 15:54 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 15:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] (duration: 09m 43s)
* 15:41 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 15:40 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:36 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]]
* 15:17 urbanecm: Delete all running periodic jobs starting with `growthexperiments-refreshlinkrecommendations-*` (to pick up new configuration; [[phab:T392944|T392944]])
* 15:08 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:06 moritzm: installing grub2 bugfix updates from Bookworm point release
* 15:04 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp6008.drmrs.wmnet
* 15:01 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:42 hnowlan: half concurrency for parsoidCachePrewarm RecordLintJob and refreshLinks in jobqueue, eqiad & codfw
* 14:35 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 14:34 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:32 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 14:20 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:07 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2013.codfw.wmnet with OS trixie
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 14:06 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 13:55 btullis@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'.
* 13:53 btullis@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'.
* 13:43 btullis@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 13:42 btullis@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 13:29 moritzm: pruned obsolete Bullseye image dispatch from the docker registry [[phab:T416452|T416452]]
* 13:28 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:26 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b7-eqiad
* 13:25 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 13:24 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:22 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 13:22 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:17 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a4-eqiad
* 13:17 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] (duration: 10m 15s)
* 13:14 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 13:11 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 13:10 sbisson@deploy1003: sbisson: Continuing with deployment
* 13:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:04 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]]
* 13:02 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on ldap-rw[1001,2001].wikimedia.org with reason: work in progress
* 12:49 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 12:48 btullis@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
* 12:46 btullis@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
* 12:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] (duration: 14m 39s)
* 12:41 ladsgroup@deploy1003: tryvix1509, ladsgroup: Continuing with deployment
* 12:35 ladsgroup@deploy1003: tryvix1509, ladsgroup: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:31 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]]
* 12:16 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 12:16 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 11:53 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] (duration: 21m 58s)
* 11:48 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 11:35 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:31 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]]
* 10:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Repooling db2196
* 10:47 moritzm: pruned obsolete Bullseye images nodejs12-slim/nodejs12-devel/nodejs14-slim/nodejs16-slim from the docker registry [[phab:T416452|T416452]]
* 10:43 moritzm: installing Bird security updates
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1260: Repooling after cloning
* 10:09 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Repooling db2196
* 10:07 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 10:06 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 09:55 moritzm: pruned obsolete Bullseye images openjdk-8-jdk/openjdk-8-jre/openjdk-11-jre/openjdk-11-jdk from the docker registry [[phab:T416452|T416452]]
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1260: Repooling after cloning
* 09:52 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 09:52 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 09:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:28 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:27 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:03 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 09:02 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 09:01 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 16 hosts with reason: upgrade ssw1-a1-eqiad
* 08:58 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 22 hosts with reason: upgrade ssw1-a1-eqiad
* 08:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 08:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 08:49 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 08:49 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 08:48 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 08:48 cmooney@cumin1004: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 08:36 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning sanitarium
* 08:30 brouberol@dns1004: END - running authdns-update
* 08:28 moritzm: pruned obsolete Bullseye image golang1.15 from the docker registry [[phab:T416452|T416452]]
* 08:28 brouberol@dns1004: START - running authdns-update
* 08:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1260: Needs to clone another host from this one
* 08:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1260: Needs to clone another host from this one
* 08:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1260.eqiad.wmnet with reason: Cloning sanitarium
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:00 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 07:40 chlod: UTC morning backport window done
* 07:37 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] (duration: 21m 36s)
* 07:32 chlod@deploy1003: chlod, hamishz: Continuing with deployment
* 07:20 chlod@deploy1003: chlod, hamishz: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 45s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1025.eqiad.wmnet with OS bookworm
== 2026-09-08 ==
* 23:51 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:48 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:39 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1313.eqiad.wmnet
* 23:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:25 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 23:19 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:19 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:15 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:05 Amir1: dropped 57 tables on db1260 ([[phab:T437278|T437278]])
* 23:03 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 23:03 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:02 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 22:57 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1313
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1313
* 22:37 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1313
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:37 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:33 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 22:33 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1313
* 22:32 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 22:32 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1313.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1306.eqiad.wmnet
* 22:27 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6008.drmrs.wmnet with OS trixie
* 22:18 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 22:03 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 22:01 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] (duration: 09m 53s)
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 21:58 Amir1: drop links tables from db2210 ([[phab:T437278|T437278]])
* 21:57 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 21:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:56 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:52 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 21:51 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]]
* 21:51 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:48 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:45 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] (duration: 05m 27s)
* 21:43 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Continuing with deployment
* 21:40 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:39 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]]
* 21:35 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1024.eqiad.wmnet with OS bookworm
* 21:33 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp6008.drmrs.wmnet with OS trixie
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1306
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1306
* 21:30 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1306
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:24 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] (duration: 09m 12s)
* 21:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 21:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1306
* 21:19 reedy@deploy1003: reedy: Continuing with deployment
* 21:19 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:19 reedy@deploy1003: reedy: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1306.eqiad.wmnet
* 21:15 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:15 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]]
* {{safesubst:SAL entry|1=21:10 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out}}
* 21:05 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1024.eqiad.wmnet with OS bookworm
* 21:05 sbassett@deploy1003: sbassett: Continuing with deployment
* 21:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 21:04 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=21:03 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out non-http(s) lice}}
* 20:59 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=20:58 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out n}}
* 20:53 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 20:50 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1305.eqiad.wmnet
* 20:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 20:34 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 20:28 sbassett@deploy1003: sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:27 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-http(s) license}}
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1023.eqiad.wmnet with OS bookworm
* {{safesubst:SAL entry|1=20:23 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-}}
* 20:15 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] (duration: 10m 16s)
* 20:13 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:10 aaron@deploy1003: aaron: Continuing with deployment
* 20:09 aaron@deploy1003: aaron: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 20:05 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:05 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]]
* 20:01 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1023.eqiad.wmnet with OS bookworm
* 19:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1305
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1305
* 19:43 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1305
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:43 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:39 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 19:39 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1305
* 19:38 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 19:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1305.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1275.eqiad.wmnet
* 19:23 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 19:17 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 18:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:53 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1275
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1275
* 18:33 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1275
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1275
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 18:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1275.eqiad.wmnet
* 18:18 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 17:43 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy1003.eqiad.wmnet
* 17:36 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet
* 17:36 cdobbins@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-ntp (exit_code=0) rolling restart_daemons on A:dnsbox
* 17:30 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy1003.eqiad.wmnet
* 17:25 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet
* 17:15 swfrench@deploy1003: Finished scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup (duration: 04m 18s)
* 17:11 Amir1: dropping links tables from db1247 (s4 replica) - ([[phab:T437278|T437278]])
* 17:10 swfrench@deploy1003: Started scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup
* 16:51 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] (duration: 10m 19s)
* 16:46 zabe@deploy1003: zabe: Continuing with deployment
* 16:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]]
* 16:29 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 jhancock@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:12 jhancock@cumin2003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 16:08 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2003.codfw.wmnet
* 15:58 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:55 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2003.codfw.wmnet
* 15:54 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:44 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2002.codfw.wmnet
* 15:19 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2002.codfw.wmnet
* 14:55 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:46 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2001.codfw.wmnet
* 14:44 topranks: shutdown et-1/1/5 on cr1-codfw to shift traffic off ssw1-a1-codfw
* 14:43 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: upgrade ssw1-a1-eqiad
* 14:34 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2001.codfw.wmnet
* 14:33 btullis@cumin1003: END (ERROR) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=97) rolling reboot on A:cephosd-codfw
* 14:30 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet
* 14:28 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet
* 14:22 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet
* 14:13 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --olderThan {{Gerrit|1788220800}} --verbose # [[phab:T437158|T437158]]
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet
* 14:12 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet
* 14:03 topranks: drain traffic from ssw1-a1-codfw before JunOS upgrade [[phab:T426197|T426197]]
* 14:02 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet
* 13:58 cgoubert@deploy1003: helmfile [staging-codfw] DONE helmfile.d/services/mw-debug: apply
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet
* 13:57 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet
* 13:56 cgoubert@deploy1003: helmfile [staging-codfw] START helmfile.d/services/mw-debug: apply
* 13:55 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:54 stran@deploy1003: mwscript-k8s job started: foreachwikiindblist checkuser-suggested-investigations extensions/CheckUser/maintenance/populateSiCaseProperties.php # [[phab:T435066|T435066]]
* 13:52 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:51 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet
* 13:50 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:46 cdobbins@cumin1003: START - Cookbook sre.dns.roll-restart-ntp rolling restart_daemons on A:dnsbox
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet
* 13:46 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet
* 13:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2012.codfw.wmnet with OS bookworm
* 13:44 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] (duration: 34m 00s)
* 13:40 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet
* 13:37 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet
* 13:35 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet
* 13:32 stran@deploy1003: stran: Continuing with deployment
* 13:29 stran@deploy1003: stran: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet
* 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:21 moritzm: installing qemu security updates
* 13:18 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet
* 13:16 ayounsi@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'email' for AS: 139628
* 13:15 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 139628
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet
* 13:13 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet
* 13:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 2519
* 13:11 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 2519
* 13:10 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 14593
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]]
* 13:09 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:08 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 14593
* 13:06 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet
* 13:06 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2012.codfw.wmnet with OS bookworm
* 13:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2012.codfw.wmnet
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2012.codfw.wmnet
* 13:05 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:04 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet
* 13:01 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet
* 12:58 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 34655
* 12:58 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 34655
* 12:56 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2012.codfw.wmnet
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'clear' for AS: 35320
* 12:55 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'clear' for AS: 35320
* 12:55 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-codfw
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-codfw
* 12:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2011.codfw.wmnet with OS bookworm
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-by27-esams
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-by27-esams
* 12:54 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-esams
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-esams
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-esams
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-esams
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-eqsin
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-eqsin
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-ulsfo
* 12:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f3-eqiad
* 12:50 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f3-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-c8-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-c8-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:44 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:43 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet
* 12:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-eqiad
* 12:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f1-eqiad
* 12:41 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f1-eqiad
* 12:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-d5-eqiad
* 12:40 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-d5-eqiad
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet
* 12:38 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet
* 12:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:32 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet
* 12:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad
* 12:11 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2012.codfw.wmnet
* 12:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2011.codfw.wmnet with OS bookworm
* 12:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2011.codfw.wmnet
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2011.codfw.wmnet
* 12:09 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:06 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] (duration: 09m 54s)
* 12:01 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2011.codfw.wmnet
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Continuing with deployment
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2010.codfw.wmnet with OS bookworm
* 11:56 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]]
* 11:46 marostegui@dns1004: END - running authdns-update
* 11:44 marostegui@dns1004: START - running authdns-update
* 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:40 Amir1: dropping unneeded tables from x4 - db1260 ([[phab:T437278|T437278]])
* 11:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:23 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2011.codfw.wmnet
* 11:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2010.codfw.wmnet with OS bookworm
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2010.codfw.wmnet
* 11:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2010.codfw.wmnet
* 11:20 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2010.codfw.wmnet
* 11:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2009.codfw.wmnet with OS bookworm
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:43 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] (duration: 10m 57s)
* 10:39 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:38 samtar@deploy1003: samtar: Continuing with deployment
* 10:37 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:37 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:36 samtar@deploy1003: samtar: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:34 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:32 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]]
* 10:30 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2009.codfw.wmnet with OS bookworm
* 10:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2009.codfw.wmnet
* 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:18 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2009.codfw.wmnet
* 10:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2008.codfw.wmnet with OS bookworm
* 10:07 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:05 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:01 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:45 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] (duration: 13m 15s)
* 09:45 ayounsi@dns1004: END - running authdns-update
* 09:44 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:43 ayounsi@dns1004: START - running authdns-update
* 09:39 zabe@deploy1003: zabe: Continuing with deployment
* 09:37 zabe@deploy1003: zabe: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2008.codfw.wmnet with OS bookworm
* 09:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:32 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]]
* 09:32 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2008.codfw.wmnet
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2008.codfw.wmnet
* 09:31 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:29 XioNoX: remove GRE tunnels eqiad-drmrs eqdfw-ulsfo
* 09:23 moritzm: installing rsync security updates
* 09:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2007.codfw.wmnet with OS bookworm
* 09:22 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2008.codfw.wmnet
* 09:11 marostegui@cumin1003: dbctl commit (dc=all): 'Make x4 and s4 RW again [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96393 and previous config saved to /var/cache/conftool/dbconfig/20260908-091121-marostegui.json
* 09:07 marostegui@cumin1003: dbctl commit (dc=all): 'Remove old s4 masters from x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96392 and previous config saved to /var/cache/conftool/dbconfig/20260908-090749-marostegui.json
* 09:05 marostegui@cumin1003: dbctl commit (dc=all): 'Set x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96391 and previous config saved to /var/cache/conftool/dbconfig/20260908-090517-marostegui.json
* 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Set s4 commons to read-only for maintenance [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96389 and previous config saved to /var/cache/conftool/dbconfig/20260908-090228-marostegui.json
* 09:02 marostegui: Starting x4 split from s4, RO time on commons needed [[phab:T404715|T404715]]
* 09:00 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 08:58 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2008.codfw.wmnet
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:43 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 32 hosts with reason: x4 split
* 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2007.codfw.wmnet with OS bookworm
* 08:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:37 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:37 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2007.codfw.wmnet
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2007.codfw.wmnet
* 08:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:34 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:29 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:26 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2007.codfw.wmnet
* 08:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2006.codfw.wmnet with OS bookworm
* 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:50 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2007.codfw.wmnet
* 07:43 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2006.codfw.wmnet with OS bookworm
* 07:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2006.codfw.wmnet
* 07:37 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:30 denisse: Add grafana-plugins 0.15 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 07:29 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:27 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 07:27 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:27 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:22 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 07:22 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 07:18 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:14 jmm@dns1004: END - running authdns-update
* 07:13 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96388 and previous config saved to /var/cache/conftool/dbconfig/20260908-071308-marostegui.json
* 07:12 jmm@dns1004: START - running authdns-update
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96387 and previous config saved to /var/cache/conftool/dbconfig/20260908-071216-marostegui.json
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96386 and previous config saved to /var/cache/conftool/dbconfig/20260908-071159-marostegui.json
* 05:07 denisse: Add grafana-plugins 0.10 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.16 (duration: 02m 27s)
* 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]] (duration: 36m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 41s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-07 ==
* 21:52 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] (duration: 11m 00s)
* 21:47 zabe@deploy1003: zabe: Continuing with deployment
* 21:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]]
* 21:37 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] (duration: 09m 34s)
* 21:33 zabe@deploy1003: zabe: Continuing with deployment
* 21:32 zabe@deploy1003: zabe: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:28 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]]
* 21:03 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] (duration: 10m 27s)
* 20:58 zabe@deploy1003: zabe: Continuing with deployment
* 20:57 zabe@deploy1003: zabe: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:52 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]]
* 20:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set weight of db1261 to zero in s4 ([[phab:T437108|T437108]])', diff saved to https://phabricator.wikimedia.org/P96385 and previous config saved to /var/cache/conftool/dbconfig/20260907-203804-ladsgroup.json
* 20:23 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] (duration: 09m 28s)
* 20:19 zabe@deploy1003: zabe: Continuing with deployment
* 20:18 zabe@deploy1003: zabe: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:14 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]]
* 20:12 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] (duration: 10m 06s)
* 20:07 zabe@deploy1003: zabe, daimona: Continuing with deployment
* 20:06 zabe@deploy1003: zabe, daimona: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:02 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]]
* 19:59 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] (duration: 11m 27s)
* 19:55 zabe@deploy1003: zabe: Continuing with deployment
* 19:52 zabe@deploy1003: zabe: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:48 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]]
* 19:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] (duration: 11m 40s)
* 19:26 zabe@deploy1003: zabe: Continuing with deployment
* 19:23 zabe@deploy1003: zabe: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:19 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]]
* 19:07 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] (duration: 14m 17s)
* 19:00 zabe@deploy1003: zabe: Continuing with deployment
* 18:57 zabe@deploy1003: zabe: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:53 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]]
* 18:33 zabe@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 18:32 zabe@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 18:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] (duration: 09m 12s)
* 18:27 zabe@deploy1003: zabe: Continuing with deployment
* 18:26 zabe@deploy1003: zabe: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:22 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]]
* 16:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2005.codfw.wmnet with OS bookworm
* 16:01 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] (duration: 10m 22s)
* 15:59 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 15:57 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 15:56 zabe@deploy1003: zabe: Continuing with deployment
* 15:55 zabe@deploy1003: zabe: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:51 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]]
* 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:47 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 15:46 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 15:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:44 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 15:44 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 15:27 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2005.codfw.wmnet with OS bookworm
* 15:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2005.codfw.wmnet
* 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:11 moritzm: installing rsync security updates
* 15:04 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1341.eqiad.wmnet
* 15:03 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2004.codfw.wmnet with OS bookworm
* 15:00 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:58 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 14:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:40 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 14:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 14:35 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:32 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:29 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:25 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 14:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1228: Repooling db1228 into s4
* 14:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2004.codfw.wmnet with OS bookworm
* 14:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Repooling after cloning
* 14:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:19 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2004.codfw.wmnet
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2004.codfw.wmnet
* 14:18 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:14 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1074.eqiad.wmnet
* 14:14 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:13 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:13 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* {{safesubst:SAL entry|1=14:11 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mul}}
* 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2004.codfw.wmnet
* 14:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1074.eqiad.wmnet
* 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1073.eqiad.wmnet
* 14:07 krinkle@deploy1003: krinkle: Continuing with deployment
* {{safesubst:SAL entry|1=14:04 krinkle@deploy1003: krinkle: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with multiple properties}}
* 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1073.eqiad.wmnet
* 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1072.eqiad.wmnet
* 14:01 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1341
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1341
* 14:01 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 14:00 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1341
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* 13:59 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* {{safesubst:SAL entry|1=13:59 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mult}}
* 13:59 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2004.codfw.wmnet
* 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2003.codfw.wmnet with OS bookworm
* 13:55 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 13:55 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1072.eqiad.wmnet
* 13:55 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudvirt1067.eqiad.wmnet
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1341
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 13:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1341.eqiad.wmnet
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2008.wikimedia.org
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2008.wikimedia.org with OS trixie
* 13:39 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1067.eqiad.wmnet
* 13:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1066.eqiad.wmnet
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:37 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1228: Repooling db1228 into s4
* 13:36 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Repooling after cloning
* 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1066.eqiad.wmnet
* 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1065.eqiad.wmnet
* 13:28 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:27 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1065.eqiad.wmnet
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:25 moritzm: installing openssh security updates
* 13:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:24 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1340.eqiad.wmnet
* 13:23 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] (duration: 10m 06s)
* 13:17 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2003.codfw.wmnet with OS bookworm
* 13:17 stran@deploy1003: stran: Continuing with deployment
* 13:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:16 stran@deploy1003: stran: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:16 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:15 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 13:12 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]]
* 13:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2003.codfw.wmnet
* 13:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2003.codfw.wmnet
* 13:03 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2008.wikimedia.org with OS trixie
* 13:03 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:02 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:02 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:01 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:00 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2003.codfw.wmnet
* 12:58 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:54 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 12:54 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2008.wikimedia.org
* 12:47 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2003.codfw.wmnet
* 12:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2002.codfw.wmnet with OS bookworm
* 12:29 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 12:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2007.wikimedia.org
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2007.wikimedia.org with OS trixie
* 12:14 elukey: moved most of the Docker Registry's prefixes to a new internal S3 backend. For any docker pull failure that worked in the past, please ping me or drop a note in [[phab:T435499|T435499]] or contact the oncall SREs
* 12:07 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2002.codfw.wmnet with OS bookworm
* 12:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:02 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2002.codfw.wmnet
* 11:54 jmm@dns1004: END - running authdns-update
* 11:52 jmm@dns1004: START - running authdns-update
* 11:46 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2002.codfw.wmnet
* 11:46 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2007.wikimedia.org with OS trixie
* 11:46 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:46 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:41 moritzm: installing bash updates from bookworm point release
* 11:39 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:39 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2007.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts ldap-replica1006.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:35 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:32 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2001.codfw.wmnet with OS bookworm
* 11:28 jmm@cumin1004: START - Cookbook sre.hosts.decommission for hosts ldap-replica1006.wikimedia.org
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 11:18 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] (duration: 14m 08s)
* 11:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 11:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet
* 11:11 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:11 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:10 zabe@deploy1003: zabe: Continuing with deployment
* 11:10 zabe@deploy1003: zabe: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 11:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet
* 11:07 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet
* 11:07 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet
* 11:06 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:06 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:04 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]]
* 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 10:58 btullis@deploy1003: Finished scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]] (duration: 35m 20s)
* 10:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:53 jmm@dns1004: END - running authdns-update
* 10:51 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:50 jmm@dns1004: START - running authdns-update
* 10:47 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=frwiki # [[phab:T436659|T436659]]
* 10:40 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=hrwiki # [[phab:T436659|T436659]]
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:39 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 10:37 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 10:33 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 10:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2001.codfw.wmnet with OS bookworm
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1340
* 10:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1340.eqiad.wmnet
* 10:26 btullis@deploy1003: Started scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]]
* 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2001.codfw.wmnet
* 10:18 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2001.codfw.wmnet
* 10:13 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw
* 10:12 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:10 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:09 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:08 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:07 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:07 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:02 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:00 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:59 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:57 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/services/miscweb: apply
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/services/miscweb: apply
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:52 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:52 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] (duration: 10m 11s)
* 09:51 aokoth@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 09:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 09:49 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 09:49 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 09:48 aokoth@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 09:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:45 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:44 aokoth@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 09:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]]
* 09:41 aokoth@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 09:38 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Repooling after switchover
* 09:23 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw
* 09:15 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] (duration: 20m 12s)
* 09:12 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 139009
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1005.wikimedia.org
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1005.wikimedia.org with OS trixie
* 09:10 moritzm: rebuild software RAID following disk replacement [[phab:T437036|T437036]]
* 09:10 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 139009
* 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1022.eqiad.wmnet with OS bookworm
* 09:08 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 09:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:03 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet
* 09:02 moritzm: installing giflib security updates
* 09:01 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:59 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 08:57 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 08:56 jmm@cumin1004: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet
* 08:56 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:54 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]]
* 08:52 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:52 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:49 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 08:48 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:42 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:41 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:40 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:40 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96355 and previous config saved to /var/cache/conftool/dbconfig/20260907-083904-marostegui.json
* 08:38 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1005.wikimedia.org with OS trixie
* 08:38 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2157 to s5 primary [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96354 and previous config saved to /var/cache/conftool/dbconfig/20260907-083825-marostegui.json
* 08:38 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 marostegui: Starting s5 codfw failover from db2213 to db2157 - [[phab:T437188|T437188]]
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2157 with weight 0 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96353 and previous config saved to /var/cache/conftool/dbconfig/20260907-083448-marostegui.json
* 08:34 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s5 [[phab:T437188|T437188]]
* 08:28 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:28 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1005.wikimedia.org
* 08:22 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:20 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1022.eqiad.wmnet with OS bookworm
* 08:03 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:02 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 08:02 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:00 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 07:57 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 07:56 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1241.eqiad.wmnet with reason: Cloning
* 07:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1241: Cloning
* 07:55 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1241: Cloning
* 07:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Cloning
* 07:54 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Cloning
* 07:51 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 07:50 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 07:47 kartik@deploy1003: Finished scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] (duration: 41m 51s)
* 07:46 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 07:45 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 07:34 kartik@deploy1003: abi, kartik: Continuing with deployment
* 07:23 kartik@deploy1003: abi, kartik: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:05 kartik@deploy1003: Started scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]]
* 06:14 moritzm: installing Chromium security updates
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 08m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-06 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 25s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-05 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 26s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-04 ==
* 22:07 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 21:48 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:42 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:30 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:42 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:40 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 19:27 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:19 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:13 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:12 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host sretest2013
* 19:11 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host sretest2013
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:11 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:07 jhancock@cumin1003: START - Cookbook sre.dns.netbox
* 18:18 inflatador: bking@clouddumps100[12] `systemctl reset-failed` to quash alerts until https://w.wiki/UBje . The systemd timer should try again tomorrow
* 17:27 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b8-eqiad
* 17:27 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b8-eqiad
* 16:37 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:33 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-b7-eqiad
* 16:33 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 16:05 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b6-eqiad
* 16:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b6-eqiad
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1339.eqiad.wmnet
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 15:50 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:49 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:47 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:43 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 15:10 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b5-eqiad
* 15:09 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b5-eqiad
* 14:46 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1045.eqiad.wmnet
* 14:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1003.eqiad.wmnet with OS trixie
* 14:38 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 14:37 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b4-eqiad
* 14:37 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b4-eqiad
* 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339
* 14:32 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339
* 14:31 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 14:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet
* 14:26 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:24 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:21 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b3-eqiad
* 14:21 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b3-eqiad
* 14:17 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:17 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:04 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1003.eqiad.wmnet with OS trixie
* 13:54 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b2-eqiad
* 13:53 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b2-eqiad
* 13:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1002.eqiad.wmnet with OS trixie
* 13:18 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a4-eqiad
* 13:12 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:12 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b1-eqiad
* 13:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b1-eqiad
* 12:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1002.eqiad.wmnet with OS trixie
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow3004.esams.wmnet with OS trixie
* 12:33 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:28 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 12:15 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 12:11 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a7-eqiad
* 12:11 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a7-eqiad
* 12:01 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow3004.esams.wmnet with OS trixie
* 11:47 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a6-eqiad
* 11:47 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a6-eqiad
* 11:36 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki2003.codfw.wmnet
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki2003.codfw.wmnet with OS trixie
* 11:19 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:13 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a5-eqiad
* 11:06 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a5-eqiad
* 10:52 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki2003.codfw.wmnet with OS trixie
* 10:50 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:50 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:44 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 10:44 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki2003.codfw.wmnet
* 10:29 btullis@deploy1003: Finished scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli (duration: 41m 14s)
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 10:00 marostegui@cumin1003: Removing db1182 from zarcillo [[phab:T434869|T434869]]
* 10:00 marostegui@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db1182.eqiad.wmnet
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:57 btullis@deploy1003: Started scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli
* 09:53 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:53 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.decommission (exit_code=1)
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1182.eqiad.wmnet
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:50 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:46 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:45 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1182 from dbctl [[phab:T434869|T434869]]', diff saved to https://phabricator.wikimedia.org/P96346 and previous config saved to /var/cache/conftool/dbconfig/20260904-094527-marostegui.json
* 09:41 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki1003.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki1003.eqiad.wmnet with OS trixie
* 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1182: Decommissioning
* 09:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1182: Decommissioning
* 09:23 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 09:17 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow5003.eqsin.wmnet with OS trixie
* 09:02 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki1003.eqiad.wmnet with OS trixie
* 09:00 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 09:00 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:55 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:55 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki1003.eqiad.wmnet
* 08:54 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:48 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:45 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:40 btullis@deploy1003: Finished scap sync-world: Trying again for [[phab:T436913|T436913]] (duration: 34m 26s)
* 08:35 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2003.codfw.wmnet
* 08:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw2001.wikimedia.org with OS trixie
* 08:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4 days, 0:00:00 on db2196.codfw.wmnet with reason: Host crashed
* 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2003.codfw.wmnet
* 08:24 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:20 elukey@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:12 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:08 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:07 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 08:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:04 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 07:57 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow5003.eqsin.wmnet with OS trixie
* 07:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:55 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw2001.wikimedia.org with OS trixie
* 07:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:51 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:50 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie
* 07:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2196: down
* 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2196: down
* 06:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:52 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:40 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:28 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:21 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 06:20 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 39s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 01:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:02 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:59 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:46 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 00:44 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:39 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:23 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
== 2026-09-03 ==
* 21:46 tsev@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T435363|T435363]]
* 21:03 eevans@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:55 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] (duration: 12m 24s)
* 20:52 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:52 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1021.eqiad.wmnet with OS bookworm
* 20:50 arlolra@deploy1003: arlolra, tsev: Continuing with deployment
* 20:46 arlolra@deploy1003: arlolra, tsev: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:44 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1047.eqiad.wmnet
* 20:42 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]]
* 20:41 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a3-eqiad
* 20:40 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a3-eqiad
* 20:39 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] (duration: 10m 23s)
* 20:34 arlolra@deploy1003: arlolra, jgiannelos: Continuing with deployment
* 20:33 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1047.eqiad.wmnet
* 20:33 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:32 arlolra@deploy1003: arlolra, jgiannelos: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:31 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1046.eqiad.wmnet
* 20:28 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]]
* 20:23 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] (duration: 13m 41s)
* 20:20 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1046.eqiad.wmnet
* 20:16 catrope@deploy1003: catrope: Continuing with deployment
* 20:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1021.eqiad.wmnet with OS bookworm
* 20:13 catrope@deploy1003: catrope: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1021.eqiad.wmnet
* 20:09 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]]
* 20:00 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1021.eqiad.wmnet
* 19:49 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 19:19 swfrench@deploy1003: Finished scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]] (duration: 02m 59s)
* 19:16 swfrench@deploy1003: Started scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]]
* 19:01 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:01 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 18:57 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:57 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:34 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:20 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 17:55 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:55 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:53 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:52 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:46 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 17:46 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 17:45 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 17:44 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 17:44 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:40 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:39 ryankemper: [WDQS] Service looks healthy again, CPU load and thread count have dropped considerably over the last hour
* 17:39 ryankemper: [[phab:T421642|T421642]] [WDQS] requestctl changes: `2026-09-03 16:23-17:33` UTC: added hard-deny pair `cache-text/wdqs_futile_sparql_sep_2026_deny(+_bots)`; extended pattern `ua/wdqs_heavy_sparql_bots_2026` and added default-scope twin `wdqs_heavy_sparql_bots_jul_2026_ratelimit_default`; added ipblock `abuse/wdqs_sparql_scanners_sep_2026` + throttle `wdqs_sparql_scanners_sep_2026_ratelimit` (needed manual `requestctl update-provenance-map`)
* 17:37 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:37 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:31 andrew@cumin2003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudcephosd1045.eqiad.wmnet
* 17:30 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:30 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:28 dancy@deploy1003: Installation of scap version "4.289.0" completed for 3 hosts
* 17:26 dancy@deploy1003: Installing scap version "4.289.0" for 3 host(s)
* 17:24 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a2-eqiad
* 17:24 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:22 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:22 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a2-eqiad
* 17:14 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 17:10 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:10 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:09 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 17:01 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:59 andrew@cumin2003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 16:58 dancy: Running scap clean-images on deploy1003
* 16:52 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:50 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:39 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:14 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs2021.codfw.wmnet,service=wdqs-main
* 16:14 ryankemper: [[phab:T430880|T430880]] Stumbled across `wdqs2021` listed as inactive, looks like it was never fully re-pooled after a data xfer. Pooled.
* 16:12 btullis@deploy1003: Finished deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2] (duration: 00m 38s)
* 16:12 btullis@deploy1003: Started deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2]
* 16:12 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:07 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs101[1-4].eqiad.wmnet
* 16:03 btullis@deploy1003: Started scap sync-world: Rebuilding to pick up new version of dump scripts in mediawiki-cli for [[phab:T436913|T436913]]
* 16:01 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]] (duration: 09m 29s)
* 15:51 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=urldownloader[12]00[56].wikimedia.org [reason: depooling urldownloader trixie nodes]
* 15:51 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]]
* 15:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:29 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:24 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:24 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:18 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:15 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:09 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:05 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 15:04 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:02 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 15:00 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader
* 14:58 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw
* 14:53 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-ulsfo
* 14:49 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 14:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2002.codfw.wmnet
* 14:45 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:39 sukhe: sudo cumin "A:cp-text" "run-puppet-agent --enable 'merging CR 1334855'": [[phab:T425441|T425441]]
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2002.codfw.wmnet
* 14:32 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:30 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:29 sukhe: sudo cumin "A:cp-text" "disable-puppet 'merging CR 1334855'" [[phab:T425441|T425441]]
* 14:27 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-ulsfo
* 14:22 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:21 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:19 arnaudb@dns1006: END - running authdns-update
* 14:18 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1074
* 14:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1074
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-codfw
* 14:17 arnaudb@dns1006: START - running authdns-update
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:16 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:16 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1073
* 14:15 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1073
* 14:12 ayounsi@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host netflow2004.codfw.wmnet with OS trixie
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:10 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:08 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:08 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:07 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] (duration: 09m 36s)
* 14:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:05 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:03 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 14:03 samtar@deploy1003: btullis, samtar: Continuing with deployment
* 14:02 samtar@deploy1003: btullis, samtar: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:01 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:58 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]]
* 13:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:56 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:55 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:52 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqiad and A:durum
* 13:52 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-codfw
* 13:51 moritzm: installing sqlite3 security updates
* 13:51 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:51 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqiad and A:durum
* 13:49 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-codfw and A:durum
* 13:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:47 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-codfw and A:durum
* 13:47 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-esams
* 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:44 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:43 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:43 ayounsi@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:42 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] (duration: 13m 50s)
* 13:41 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:40 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1072
* 13:40 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:39 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-esams and A:durum
* 13:38 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1072
* 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:38 samtar@deploy1003: samtar, thiemowmde: Continuing with deployment
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-esams and A:durum
* 13:37 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:33 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:33 samtar@deploy1003: samtar, thiemowmde: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:32 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqsin and A:durum
* 13:31 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqsin and A:durum
* 13:28 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]]
* 13:28 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 13:27 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:24 ayounsi@cumin1004: START - Cookbook sre.hosts.reimage for host netflow2004.codfw.wmnet with OS trixie
* 13:24 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:22 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 13:22 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-esams
* 13:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:15 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:15 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 moritzm: installing bash updates from trixie point release
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 13:14 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1067
* 13:13 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1067
* 13:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2901: Test
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:09 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:09 moritzm: installing libxslt bugfix updates from Trixie point release
* 13:08 jelto@dns1004: END - running authdns-update
* 13:06 jelto@dns1004: START - running authdns-update
* 13:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 13:04 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:04 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 13:00 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:59 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1066
* 12:59 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1066
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:58 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:54 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 12:53 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2901: Test
* 12:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2901: Test
* 12:52 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'sync'.
* 12:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 12:50 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2901: Test
* 12:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:50 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 12:42 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:41 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-magru and A:durum
* 12:38 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-magru and A:durum
* 12:34 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 12:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1282: Pooling db1282 into s6
* 12:31 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 12:25 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:21 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:19 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:15 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_ulsfo
* 12:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1020.eqiad.wmnet with OS bookworm
* 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: db2207 repool
* 12:07 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_ulsfo
* 12:04 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 11:58 kart_: cxserver: Use urldownloader LVS endpoint ([[phab:T429175|T429175]])
* 11:57 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply
* 11:56 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply
* 11:56 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_magru
* 11:55 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply
* 11:55 moritzm: installing rsync security updates
* 11:55 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply
* 11:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:52 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply
* 11:51 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply
* 11:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1282: Pooling db1282 into s6
* 11:45 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1282 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96328 and previous config saved to /var/cache/conftool/dbconfig/20260903-114526-marostegui.json
* 11:43 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqiad
* 11:35 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqiad
* 11:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1020.eqiad.wmnet with OS bookworm
* 11:26 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 11:24 cgoubert@deploy1003: Finished scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter (duration: 12m 01s)
* 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: db2207 repool
* 11:22 cgoubert@deploy1003: cgoubert: Continuing with deployment
* 11:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqsin
* 11:17 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_esams
* 11:15 cgoubert@deploy1003: cgoubert: mediawiki: enable forward of fatal metrics to statsd exporter synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:14 cgoubert@deploy1003: Started scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter
* 11:10 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_esams
* 11:09 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 11:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1019.eqiad.wmnet with OS bookworm
* 11:01 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_drmrs
* 10:59 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_codfw
* 10:52 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_codfw
* 10:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:41 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 10:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 10:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow2005.codfw.wmnet
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2005.codfw.wmnet with OS trixie
* 10:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:17 btullis@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync
* 10:17 btullis@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync
* 10:16 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1019.eqiad.wmnet with OS bookworm
* 10:15 btullis@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: sync
* 10:15 btullis@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: sync
* 10:12 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s8
* 10:11 marostegui: Move s8 sanitarium from db1167 to db1281 [[phab:T434778|T434778]]
* 10:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 10:03 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 09:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:55 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2003.codfw.wmnet with OS trixie
* 09:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:43 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow2005.codfw.wmnet with OS trixie
* 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:42 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1018.eqiad.wmnet with OS bookworm
* 09:41 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_ulsfo
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:39 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2003.codfw.wmnet with reason: host reimage
* 09:39 hnowlan: fixed currently oncall pane in klaxon
* 09:38 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:38 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434751|T434751]]
* 09:37 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:37 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow2005.codfw.wmnet
* 09:35 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434775|T434775]]
* 09:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s7
* 09:34 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:33 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_ulsfo
* 09:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqiad
* 09:30 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] (duration: 09m 30s)
* 09:27 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:25 zabe@deploy1003: zabe: Continuing with deployment
* 09:25 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqiad
* 09:25 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 09:25 zabe@deploy1003: zabe: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1174 from dbctl [[phab:T436904|T436904]]', diff saved to https://phabricator.wikimedia.org/P96323 and previous config saved to /var/cache/conftool/dbconfig/20260903-092448-marostegui.json
* 09:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:21 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]]
* 09:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqsin
* 09:17 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:15 topranks: put traffic on Lumen codfw<->eqiad link as it is stable [[phab:T435810|T435810]]
* 09:14 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_esams
* 09:09 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:06 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_esams
* 09:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 09:03 marostegui: Move s6 sanitarium from db1165 to db1279 [[phab:T434775|T434775]]
* 09:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1018.eqiad.wmnet with OS bookworm
* 08:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s6
* 08:57 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_drmrs
* 08:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_codfw
* 08:55 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw
* 08:49 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_codfw
* 08:49 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:45 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 08:45 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:42 marostegui: Move s5 sanitarium from db1161 to db1275 [[phab:T434776|T434776]]
* 08:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 24 hosts with reason: Changing sanitarium master in s5
* 08:38 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_magru
* 08:37 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:37 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server (duration: 01m 21s)
* 08:37 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:36 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server
* 08:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:33 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server (duration: 01m 28s)
* 08:32 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server
* 08:27 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 08:11 moritzm: uploaded wmf-laptop 1.0.7 to apt.wikimedia.org
* 08:03 marostegui: Move s2 sanitarium from db1156 to db1271 [[phab:T434287|T434287]]
* 07:59 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2001.codfw.wmnet
* 07:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 23 hosts with reason: Changing sanitarium master in s2
* 07:49 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2001.codfw.wmnet
* 07:39 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:29 chlod: UTC morning backport window done
* 07:27 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] (duration: 11m 54s)
* 07:22 chlod@deploy1003: chlod, tryvix1509: Continuing with deployment
* 07:22 XioNoX: push pfw policies - [[phab:T436729|T436729]]
* 07:20 chlod@deploy1003: chlod, tryvix1509: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]]
* 07:15 marostegui: Power off db1228 for maintenance
* 07:13 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1228.eqiad.wmnet with reason: Onsite maintenance
* 07:01 arnaudb@dns1006: END - running authdns-update
* 06:58 arnaudb@dns1006: START - running authdns-update
* 06:54 jmm@cumin2003: END (PASS) - Cookbook sre.wdqs.restart-nginx-envoy (exit_code=0) rolling restart_daemons on A:wcqs-public
* 06:52 jmm@cumin2003: START - Cookbook sre.wdqs.restart-nginx-envoy rolling restart_daemons on A:wcqs-public
* 06:46 moritzm: installing libxml2 security updates
* 06:27 hashar: Upgrading CI Jenkins on contint1003 # [[phab:T436812|T436812]]
* 06:11 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet
* 06:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts2002.codfw.wmnet
* 06:00 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet
* 05:56 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts2002.codfw.wmnet
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 48s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:16 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1017.eqiad.wmnet with OS bookworm
== 2026-09-02 ==
* 23:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] (duration: 10m 21s)
* 23:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:50 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 23:49 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]]
* 23:38 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:38 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host aqs1017.eqiad.wmnet with OS bookworm
* 22:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] (duration: 11m 03s)
* 22:32 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 22:30 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:26 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]]
* 22:05 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] (duration: 14m 14s)
* 21:59 krinkle@deploy1003: krinkle: Continuing with deployment
* 21:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:50 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]]
* 21:44 inflatador: bking@apt1002 sudo -E private_reprepro --ignore=wrongdistribution -C matomo_plugins include bookworm-wikimedia-private matomo-plugin-customreports_5.5.0-1_amd64.changes [[phab:T431608|T431608]]
* 21:40 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] (duration: 09m 48s)
* 21:35 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:34 jforrester@deploy1003: jforrester: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:34 inflatador: bking@apt1002 sudo -E reprepro -C main include bookworm-wikimedia matomo-plugin-marketingcampaignsreporting_5.2.2-3_amd64.changes [[phab:T431608|T431608]]
* 21:30 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]]
* 21:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 21:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:03 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 20:50 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 20:27 dancy@deploy1003: Finished scap sync-world: testing (duration: 09m 21s)
* 20:18 dancy@deploy1003: Started scap sync-world: testing
* 20:18 dancy@deploy1003: Installation of scap version "4.288.0" completed for 3 hosts
* 20:16 dancy@deploy1003: Installing scap version "4.288.0" for 3 host(s)
* 19:57 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] (duration: 64m 27s)
* 19:55 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 18:57 jforrester@deploy1003: jforrester: Continuing with deployment
* 18:57 jforrester@deploy1003: jforrester: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:55 swfrench-wmf: deleted pods coredns-85b4f68d95-pk5sn coredns-85b4f68d95-22ddb coredns-85b4f68d95-49k5p in eqiad due to intermittent upstream resolution health check failures correlated with high DNS resolution latency
* 18:53 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]]
* 18:31 sukhe@dns1004: END - running authdns-update
* 18:28 sukhe@dns1004: START - running authdns-update
* 18:26 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:26 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:17 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:16 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 18:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:14 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:14 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqiad
* 18:14 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:49 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqiad
* 17:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:45 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqsin
* 17:38 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:38 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:20 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqsin
* 17:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 17:00 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 16:59 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:48 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 16:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:41 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 16:41 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 16:41 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:38 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:34 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 16:29 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_drmrs
* 16:24 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 16:14 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_drmrs
* 16:11 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-unlock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 16:10 root@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]] (duration: 48m 09s)
* 16:10 root@deploy1003: Forcefully removing global lock: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 16:10 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-unlock-scap for datacenter switchover from codfw to eqiad
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:59 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:58 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters for datacenter switchover from codfw to eqiad
* 15:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 15:58 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-restore-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-restore-ttl for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-start-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 root@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-start-maintenance for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-jobrunner: sync
* 15:56 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-jobrunner: sync
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.07-set-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period ends at: 2026-09-02 15:56:13.434320
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.07-set-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.03-set-db-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.03-set-db-readonly for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.02-set-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period starts at: 2026-09-02 15:53:47.690918
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.02-set-readonly for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.01-stop-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.01-stop-maintenance for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-reduce-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:47 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-reduce-ttl for datacenter switchover from codfw to eqiad
* 15:46 slyngshede@cumin1003: END (ERROR) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=97) for datacenter switchover from codfw to eqiad
* 15:45 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 15:44 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 15:42 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service
* 15:39 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:38 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:31 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:28 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:27 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 15:27 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from codfw to eqiad
* 15:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-lock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:22 root@deploy1003: Locking from deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-lock-scap for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks for datacenter switchover from codfw to eqiad
* 15:17 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service
* 15:15 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_magru
* 15:11 moritzm: import jenkins 2.568.3 to thirdparty/jenkins for trixie-wikimedia
* 14:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 14:44 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_magru
* 14:32 moritzm: installing pdns-recursor security updates
* 14:27 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:27 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:20 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:15 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 14:12 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 14:09 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:09 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:09 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:08 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:06 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:44 moritzm: bounce tcpircbot-logmsgbot/tcpircbot-logmsgbot_cloud on alert1002 to allow cumin1004 [[phab:T427897|T427897]]
* 13:36 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] (duration: 09m 52s)
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Continuing with deployment
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:26 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]]
* 13:25 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 13:24 moritzm: installing wireshark security updates
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 13:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:04 moritzm: import librsvg 2.60.0+dfsg-1+wmf13u1 to component/thumbor for trixie-wikimedia [[phab:T436505|T436505]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 12:44 atsuko@dns1004: END - running authdns-update
* 12:41 atsuko@dns1004: START - running authdns-update
* 12:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:35 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] (duration: 12m 50s)
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:24 dreamyjazz@deploy1003: dreamyjazz, btullis: Continuing with deployment
* 12:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:22 dreamyjazz@deploy1003: dreamyjazz, btullis: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:20 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:17 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]]
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 11:31 marostegui@cumin1003: Removing db1172 from zarcillo [[phab:T436763|T436763]]
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1172.eqiad.wmnet
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:30 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:26 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:26 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:25 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 11:25 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:24 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:20 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1172.eqiad.wmnet
* 11:20 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 11:12 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:11 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:10 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:09 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:05 slyngshede@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.*
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 11:05 slyngshede@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 11:03 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply
* 11:03 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
* 10:52 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:49 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:48 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Pool back db1242
* 10:45 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow4003.ulsfo.wmnet with OS trixie
* 10:31 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1172 from dbctl [[phab:T436763|T436763]]', diff saved to https://phabricator.wikimedia.org/P96318 and previous config saved to /var/cache/conftool/dbconfig/20260902-103152-marostegui.json
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 blake@deploy1003: Finished scap sync-world: non-build deployment for [[phab:T417800|T417800]] (duration: 05m 37s)
* 10:06 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:05 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:04 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:03 blake@deploy1003: Started scap sync-world: non-build deployment for [[phab:T417800|T417800]]
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 10:00 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:00 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:59 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:58 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:57 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:56 jmm@dns1004: END - running authdns-update
* 09:56 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:55 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:54 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:53 jmm@dns1004: START - running authdns-update
* 09:51 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:47 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1228 to dbctl [[phab:T435892|T435892]]', diff saved to https://phabricator.wikimedia.org/P96313 and previous config saved to /var/cache/conftool/dbconfig/20260902-094713-marostegui.json
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:42 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow4003.ulsfo.wmnet with OS trixie
* 09:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow7002.magru.wmnet with OS trixie
* 09:31 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:30 moritzm: installing openjdk-21 security updates
* 09:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:22 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:17 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:17 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:16 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:10 moritzm: installing openjdk-8 security updates
* 09:09 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:08 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:46 tappof: bump space for prometheus k8s-dse in eqiad
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:39 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow7002.magru.wmnet with OS trixie
* 08:36 moritzm: installing libgraphite2 security updates
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:23 Msz2001: UTC morning backport window done
* 08:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] (duration: 14m 36s)
* 08:19 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:19 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet
* 08:14 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:14 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet
* 08:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet
* 08:08 fabfur@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cp5022.eqsin.wmnet with reason: investigating
* 08:07 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]]
* 08:07 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp5022.*
* 08:07 fabfur: depooling and silencing cp5022 ([[phab:T414411|T414411]])
* 08:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet
* 08:03 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:03 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* {{safesubst:SAL entry|1=08:03 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)}}
* 07:49 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 07:49 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]] synced to the
* 07:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1004.eqiad.wmnet
* 07:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1004.eqiad.wmnet
* 07:30 jmm@dns1004: END - running authdns-update
* {{safesubst:SAL entry|1=07:27 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]}}
* 07:27 jmm@dns1004: START - running authdns-update
* 07:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] (duration: 16m 04s)
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1242: Cloning db1228
* 07:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1242: Cloning db1228
* 07:18 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1228,1242].eqiad.wmnet with reason: db1242 needs to clone db1228
* 07:17 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Continuing with deployment
* 07:12 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1228.eqiad.wmnet with OS trixie
* 07:10 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be veri
* 07:06 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]]
* 06:43 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: build: Updating npm dependencies (duration: 00m 13s)
* 06:43 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: build: Updating npm dependencies
* 06:39 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:32 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:23 slyngshede@dns1004: END - running authdns-update
* 06:21 marostegui: Drop cu* tables from s3 bswiktionary [[phab:T435965|T435965]]
* 06:20 slyngshede@dns1004: START - running authdns-update
* 06:18 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1228.eqiad.wmnet with OS trixie
* 06:13 XioNoX: re-enable magru cr1/asw1-b3 link - [[phab:T436675|T436675]]
* 05:06 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] (duration: 04m 42s)
* 05:04 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:03 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:01 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 05:01 tstarling@deploy1003: Scap cancelled without rolling back.
* 04:53 tstarling@deploy1003: tstarling: Continuing with deployment
* 04:29 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 04:25 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 43s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-01 ==
* 21:59 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] (duration: 18m 05s)
* 21:52 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:47 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]]
* 21:38 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] (duration: 23m 55s)
* 21:28 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:20 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:14 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]]
* 20:47 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1016.eqiad.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:24 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:11 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 20:11 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:57 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:57 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:56 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:54 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:52 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:47 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:45 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:42 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:41 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:40 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:39 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:32 jhancock@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudcephosd1055
* 19:32 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:30 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.* [reason: update IP addrs]
* 19:30 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 19:30 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:23 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:22 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:16 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:13 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1056
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1056
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1055
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:12 jclark@cumin1003: END (ERROR) - Cookbook sre.network.configure-switch-interfaces (exit_code=97) for host cloudcephosd1054
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1054
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:11 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:06 jclark@cumin1003: START - Cookbook sre.dns.netbox
* 19:06 sukhe@dns1004: END - running authdns-update
* 19:03 sukhe@dns1004: START - running authdns-update
* 18:14 dancy@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 17:04 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 17:00 dancy@deploy1003: Finished scap sync-world: testing (duration: 08m 07s)
* 16:52 dancy@deploy1003: Started scap sync-world: testing
* 16:48 dancy@deploy1003: sync-world aborted: testing (duration: 00m 05s)
* 16:48 dancy@deploy1003: Started scap sync-world: testing
* 16:47 dancy@deploy1003: Installation of scap version "4.287.0" completed for 156 hosts
* 16:42 dancy@deploy1003: Installing scap version "4.287.0" for 156 host(s)
* 16:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 16:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 15:51 moritzm: installing mesa security updates
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts phab1004.eqiad.wmnet
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:27 aokoth@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:20 aokoth@cumin1003: START - Cookbook sre.dns.netbox
* 15:14 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595] (duration: 05m 32s)
* 15:14 aokoth@cumin1003: START - Cookbook sre.hosts.decommission for hosts phab1004.eqiad.wmnet
* 15:11 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on phab1004.eqiad.wmnet with reason: Decom
* 15:09 joal@deploy1003: Started deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595]
* 15:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config: apply
* 15:08 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config: apply
* 14:55 hashar: Restarted Jenkins on releases1003
* 14:51 hashar: Restarted CI Jenkins on contint1003
* 14:48 hashar: Restarting Gerrit primary on gerrit2003
* 14:45 hashar: Restarted Gerrit on gerrit1003 and gerrit2002
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 14:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 14:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 14:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:24 moritzm: installing curl security updates
* 14:24 jmm@dns1004: END - running authdns-update
* 14:23 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser (duration: 00m 14s)
* 14:23 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:21 jmm@dns1004: START - running authdns-update
* 14:21 jmm@dns1004: END - running authdns-update
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 14:19 jmm@dns1004: START - running authdns-update
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet
* 14:14 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:14 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:13 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595] (duration: 07m 26s)
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Test
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Test
* 14:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet
* 14:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Test
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Test
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 14:05 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595]
* 14:05 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] (duration: 37m 37s)
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 hashar: Removed openjdk-17 packages from contint1002/contint2002 following relocation of CI Jenkins to contint1003/contint2003 # [[phab:T418521|T418521]]
* 14:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:02 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:02 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:00 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595] (duration: 00m 45s)
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 13:59 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595]
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:58 ladsgroup@dns1004: END - running authdns-update
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:57 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 ladsgroup@dns1004: START - running authdns-update
* 13:56 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:56 ladsgroup@dns1004: END - running authdns-update
* 13:55 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:53 ladsgroup@dns1004: START - running authdns-update
* 13:50 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:49 kharlan@deploy1003: kharlan: Continuing with deployment
* 13:48 kharlan@deploy1003: kharlan: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2004.wikimedia.org
* 13:42 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:41 jmm@dns1004: END - running authdns-update
* 13:40 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2004.wikimedia.org
* 13:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2003.wikimedia.org
* 13:38 jmm@dns1004: START - running authdns-update
* 13:34 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2003.wikimedia.org
* 13:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1004.wikimedia.org
* 13:26 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]]
* 13:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:25 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:24 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1004.wikimedia.org
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:23 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] (duration: 20m 24s)
* 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1003.wikimedia.org
* 13:20 fnegri@deploy1003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply
* 13:19 fnegri@deploy1003: helmfile [eqiad] START helmfile.d/services/toolhub: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:16 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1003.wikimedia.org
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:15 moritzm: bump urldownloader[12]00[34] to 8G RAM [[phab:T429175|T429175]]
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:15 fnegri@deploy1003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 fnegri@deploy1003: helmfile [codfw] START helmfile.d/services/toolhub: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:11 fnegri@deploy1003: helmfile [staging] DONE helmfile.d/services/toolhub: apply
* 13:11 aude@deploy1003: aude: Continuing with deployment
* 13:10 fnegri@deploy1003: helmfile [staging] START helmfile.d/services/toolhub: apply
* 13:07 aude@deploy1003: aude: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2004.wikimedia.org with OS bookworm
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow1004.eqiad.wmnet
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1004.eqiad.wmnet with OS trixie
* 12:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:47 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:41 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] (duration: 16m 25s)
* 12:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 12:34 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 12:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 12:33 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 12:31 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:29 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:25 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]]
* 12:22 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1004.eqiad.wmnet with OS trixie
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:20 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 12:20 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 12:16 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow1004.eqiad.wmnet
* 12:15 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2004.wikimedia.org with OS bookworm
* 12:14 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 12:07 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2003.wikimedia.org with OS bookworm
* 11:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:43 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:32 jmm@cumin2003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling restart_daemons on A:kafka-test-eqiad
* 11:26 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2003.wikimedia.org with OS bookworm
* 11:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1004.wikimedia.org with OS bookworm
* 11:12 moritzm: installing openjdk-21 security updates
* 11:12 jmm@cumin2003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling restart_daemons on A:kafka-test-eqiad
* 10:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:53 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:49 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2902: Pool back db2902
* 10:45 moritzm: installing Python 3.11 security updates
* 10:37 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1004.wikimedia.org with OS bookworm
* 10:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5022.eqsin.wmnet with OS trixie
* 10:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp5022.eqsin.wmnet on all recursors
* 10:36 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache cp5022.eqsin.wmnet on all recursors
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:30 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:04 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:04 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2902: Pool back db2902
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:03 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2902: test
* 10:03 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 10:01 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 09:58 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:50 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/echoserver: apply
* 09:49 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/echoserver: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Change db1176 and db2230's weight, test-s4 masters, to 0 to mimic the rest of production [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96292 and previous config saved to /var/cache/conftool/dbconfig/20260901-092444-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96291 and previous config saved to /var/cache/conftool/dbconfig/20260901-090233-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96290 and previous config saved to /var/cache/conftool/dbconfig/20260901-090158-marostegui.json
* 09:01 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96289 and previous config saved to /var/cache/conftool/dbconfig/20260901-090121-marostegui.json
* 09:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:56 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:50 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1003.wikimedia.org with OS bookworm
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1074.eqiad.wmnet
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:46 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:43 marostegui@cumin1003: dbctl commit (dc=all): 'Test repool db2902', diff saved to https://phabricator.wikimedia.org/P96288 and previous config saved to /var/cache/conftool/dbconfig/20260901-084317-marostegui.json
* 08:42 marostegui@cumin1003: dbctl commit (dc=all): 'Test depool db2902', diff saved to https://phabricator.wikimedia.org/P96287 and previous config saved to /var/cache/conftool/dbconfig/20260901-084249-marostegui.json
* 08:39 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:37 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2902: test
* 08:36 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96285 and previous config saved to /var/cache/conftool/dbconfig/20260901-083557-marostegui.json
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96284 and previous config saved to /var/cache/conftool/dbconfig/20260901-083527-marostegui.json
* 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96283 and previous config saved to /var/cache/conftool/dbconfig/20260901-083432-marostegui.json
* 08:32 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1074.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1073.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:27 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:24 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 08:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022']
* 08:16 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1073.eqiad.wmnet
* 08:15 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1003.wikimedia.org with OS bookworm
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1072.eqiad.wmnet
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:14 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:10 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:05 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1072.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1067.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:00 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:56 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:54 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022']
* 07:53 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 07:52 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie
* 07:52 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 07:50 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1067.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1066.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:47 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 07:41 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:34 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1066.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1065.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:32 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:27 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:23 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1065.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1075.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:15 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 06:49 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:45 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1075.eqiad.wmnet
* 06:29 moritzm: installing Java 17 security updates
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.15 (duration: 02m 25s)
* 03:50 denisse@deploy1003: Finished deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2 (duration: 00m 19s)
* 03:50 denisse@deploy1003: Started deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2
* 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]] (duration: 37m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:30 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 00:29 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
== See also==
{{:Server Admin Log/Archives}}
sikfrvcri7e7osec6bdan66c7poelkp
User:SSapaty (WMF)/PTD/migrate build service job to ptd
2
460699
2458709
2026-09-20T06:22:00Z
SSapaty (WMF)
28955
format ascii diagram and code using code sample template
2458709
wikitext
text/x-wiki
= Migrating an existing Build Service job to Push-to-Deploy =
{{Note|Push-to-Deploy is currently a beta feature of Toolforge.}}
This tutorial explains how to migrate an existing Toolforge job that uses an image built with the [[Help:Toolforge/Build Service|Build Service]] to the Components Service, and then configure [[Help:Toolforge/Deploy your tool|Push-to-Deploy]] so that future Git pushes automatically trigger deployments.
If you already use the Build Service together with the Jobs Framework, most of the information needed to create the Components configuration can be generated automatically from your existing job.
This tutorial uses a scheduled job named <code>recent-changes</code> as an example.
The migration has two main stages:
# Migrate the existing Build Service job to the Components Service and verify that the component works.
# Configure your Git repository to trigger Components deployments automatically when you push changes.
The second stage enables Push-to-Deploy.
== Before you begin ==
This tutorial assumes that:
* you already have a working Toolforge job;
* the job uses an image built with the Toolforge Build Service;
* the source code is stored in a publicly accessible Git repository supported by the Build Service;
* you have access to the Toolforge tool account;
* you can commit and push changes to the source repository.
Become your tool before running Toolforge commands in this tutorial:
{{Codesample|lang=shell-session|scheme=light|code=
$ become my-tool
}}
Replace <code>my-tool</code> with the name of your Toolforge tool.
== Inspect your existing job ==
First, inspect the current Jobs Framework configuration:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge jobs list -o long
}}
For example, the <code>recent-changes</code> job looks similar to this:
{{Codesample|lang=text|scheme=light|code=
Job name: recent-changes
Command: recent-changes
Job type: scheduled: */5 * * * *
Image: tool-komla-test2/recent-changes:latest@sha256:...
File log: no
Emails: onfailure
Resources: mem: default, cpu: 0.5
Mounts: none
Retry: yes: 1 time(s)
Timeout: yes: 120s
}}
The important detail is the image:
{{Codesample|lang=text|scheme=light|code=
tool-komla-test2/recent-changes:latest
}}
An image whose name is under your tool's Build Service namespace, such as <code>tool-my-tool/my-image</code>, indicates that the job is using a Build Service image.
You can also inspect previous builds:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge build list
}}
For the example job, the relevant build points to:
{{Codesample|lang=text|scheme=light|code=
Source repository: https://gitlab.wikimedia.org/repos/cloud/wmcs/ptd-samples/tutorial-1
Image name: recent-changes
Git ref: main
}}
== Generate a Components configuration ==
The Components Service can generate an example configuration from your existing Jobs Framework jobs.
Run:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components config generate
}}
{{Warning|The generated configuration is an example. Review and validate it before deploying it.}}
For the <code>recent-changes</code> job, the command generates a configuration similar to:
{{Codesample|lang=yaml|scheme=light|name=toolforge.yaml|code=
config_version: v1beta1
components:
recent-changes:
build:
repository: https://gitlab.wikimedia.org/repos/cloud/wmcs/ptd-samples/tutorial-1
ref: main
run:
command: recent-changes
cpu: '0.5'
emails: onfailure
filelog: false
memory: 0.5Gi
mount: none
retry: 1
schedule: '*/5 * * * *'
timeout: 120
component_type: scheduled
}}
The generated configuration translates the existing Jobs Framework configuration into a Components Service configuration.
For example:
{| class="wikitable"
! Existing job setting
! Components configuration setting
|-
| Job name
| Component name
|-
| Build Service source repository
| <code>build.repository</code>
|-
| Git branch or ref
| <code>build.ref</code>
|-
| Job command
| <code>run.command</code>
|-
| CPU
| <code>run.cpu</code>
|-
| Memory
| <code>run.memory</code>
|-
| Email notification policy
| <code>run.emails</code>
|-
| File logging
| <code>run.filelog</code>
|-
| Filesystem mount
| <code>run.mount</code>
|-
| Retry count
| <code>run.retry</code>
|-
| Job timeout
| <code>run.timeout</code>
|-
| Cron schedule
| <code>run.schedule</code>
|-
| Job type
| <code>component_type</code>
|}
== Review the generated configuration ==
Do not deploy the generated configuration without reviewing it.
Check that at least the following values match your existing deployment:
* source repository;
* Git branch or ref;
* command;
* component type;
* CPU and memory;
* schedule, for scheduled jobs;
* retry and timeout settings;
* email notification settings;
* filesystem mount settings.
You should also verify that the generated component represents the job you intend to migrate.
For example:
{{Codesample|lang=yaml|scheme=light|code=
components:
recent-changes:
}}
corresponds to the existing Jobs Framework job named <code>recent-changes</code>.
If your tool contains several jobs, <code>toolforge components config generate</code> may generate configuration for multiple components. Review each generated component before continuing.
== Add the configuration to your source repository ==
The Components configuration can be hosted anywhere that is publicly accessible.
For this tutorial, save the generated configuration as <code>toolforge.yaml</code> in the same Git repository as the source code. Keeping the source code and deployment configuration together makes it easier to track changes to both in Git.
You can use any publicly accessible Git repository. For Toolforge projects, we recommend using a repository in the Wikimedia GitLab <code>toolforge-repos</code> namespace.
For example:
{{Codesample|lang=text|scheme=light|code=
my-project/
├── Procfile
├── requirements.txt
├── source-files...
└── toolforge.yaml
}}
Commit and push the file:
{{Codesample|lang=shell-session|scheme=light|code=
$ git add toolforge.yaml
$ git commit -m "Add Toolforge Components configuration"
$ git push
}}
The repository and ref in <code>toolforge.yaml</code> should point to the source you want Toolforge to build.
== Create the Components configuration ==
Once <code>toolforge.yaml</code> is publicly accessible, create the Components Service configuration from the copy stored in the repository.
For example:
{{Codesample|lang=shell-session|scheme=light|code=
$ curl 'https://gitlab.wikimedia.org/toolforge-repos/my-tool/-/raw/main/toolforge.yaml?ref_type=heads' {{!}} toolforge components config create
}}
Replace the URL with the public raw URL for your <code>toolforge.yaml</code> file.
This avoids having to clone the repository or download <code>toolforge.yaml</code> on the Toolforge bastion.
Inspect the stored configuration:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components config show
}}
Review the output and make sure it matches the configuration you committed.
== Test the migration with a manual deployment ==
Before configuring Push-to-Deploy, create a deployment manually to verify that the Components Service can successfully manage the workload.
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components deployment create
}}
{{Note|This is a manual Components Service deployment. Running <code>toolforge components deployment create</code> yourself is not Push-to-Deploy. Push-to-Deploy will be configured later in this tutorial so that your CI runner creates deployments automatically after you push changes to Git.}}
The Components Service will use the configuration to build and deploy the component.
You can optionally attach a description:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components deployment create --description "Test recent-changes migration"
}}
Normally, the Components Service can reuse an existing build when the source has not changed.
To force a new build, use:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components deployment create --force-build
}}
To force the component to run again even when its configuration has not changed, use:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components deployment create --force-run
}}
These options are normally unnecessary for the initial migration.
== Verify the migrated component ==
Inspect your Components deployments:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components deployment list
}}
You can also inspect the current Components configuration:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components config show
}}
For a scheduled component, verify that the resulting job exists:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge jobs list -o long
}}
Compare the resulting job with the configuration you reviewed earlier.
For the example, you would expect to see a job similar to:
{{Codesample|lang=text|scheme=light|code=
Job name: recent-changes
Command: recent-changes
Job type: scheduled: */5 * * * *
}}
Also verify that the job behaves as expected when its next scheduled execution occurs.
At this point, you have migrated the workload to the Components Service and verified that it works. The next step is to configure your Git repository so that future deployments are triggered automatically.
== Configure Push-to-Deploy ==
Push-to-Deploy allows a CI runner to trigger a Components Service deployment after changes are pushed to your Git repository.
Follow the instructions in [[Help:Toolforge/Deploy your tool#Triggering a deployment from your CI runner|Triggering a deployment from your CI runner]] to configure your repository.
The setup gives the CI runner the credentials it needs to request a Components deployment for your tool.
After the CI configuration has been added to the repository, future deployments can follow this workflow:
{{Codesample|lang=text|scheme=light|code=
Git push
↓
CI runner
↓
Components Service deployment
↓
Toolforge workload
}}
You no longer need to SSH to Toolforge and manually run <code>toolforge components deployment create</code> for normal deployments.
== Test Push-to-Deploy ==
After configuring your CI runner, make a small change to <code>toolforge.yaml</code>.
For example, change the schedule:
{{Codesample|lang=yaml|scheme=light|name=toolforge.yaml|code=
run:
schedule: '0 * * * *'
}}
Commit and push the change:
{{Codesample|lang=shell-session|scheme=light|code=
$ git add toolforge.yaml
$ git commit -m "Change recent-changes schedule"
$ git push
}}
Do not manually create a deployment.
The push should trigger your CI runner, which creates a Components Service deployment automatically.
Check your repository's CI pipeline and verify that it completes successfully.
You can then inspect the deployment from Toolforge:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components deployment list
}}
For the scheduled example, verify that the new schedule has been applied:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge jobs list -o long
}}
If the deployment was triggered automatically and the new configuration was applied, Push-to-Deploy is working.
== Avoid independently managing the migrated job ==
After migrating a scheduled or continuous job, do not maintain a second independently managed copy of the same workload.
The component created by the Components Service may appear when you inspect workloads using Jobs Framework commands. This does not mean that you should continue managing that workload separately with <code>toolforge jobs</code> commands.
After migration, make changes to the component through <code>toolforge.yaml</code> and your source repository.
== Make future changes through Git ==
Once Push-to-Deploy is configured, make changes to your source code or <code>toolforge.yaml</code>, commit them, and push them to the configured Git branch.
For example:
{{Codesample|lang=shell-session|scheme=light|code=
$ git add .
$ git commit -m "Update recent-changes"
$ git push
}}
The CI runner triggers the deployment. You do not normally need to manually rebuild the Build Service image, update the Jobs Framework job, or create a Components deployment from the Toolforge bastion.
This allows the Git repository to track the source code and, when you choose to keep them together, the Toolforge deployment configuration.
== What changed? ==
Before migration, the deployment workflow is approximately:
{{Codesample|lang=text|scheme=light|code=
Git repository
|
v
Build Service
|
v
Build Service image
|
v
Jobs Framework job
}}
The maintainer manages the build and the job separately.
After migrating to the Components Service and configuring Push-to-Deploy, the workflow becomes:
{{Codesample|lang=text|scheme=light|code=
Git repository
├── source code
└── toolforge.yaml
↓
Git push
↓
CI runner
↓
Components Service deployment
├── build
└── run
↓
Toolforge workload
}}
The Components configuration describes how the workload should be built and run, while Push-to-Deploy allows changes pushed to Git to trigger deployment automatically.
== Troubleshooting ==
=== Check the generated configuration ===
If the generated configuration does not look correct, do not deploy it immediately.
Compare it with:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge jobs list -o long
$ toolforge build list
}}
Correct <code>toolforge.yaml</code> before creating the Components configuration.
=== Check the Components configuration ===
Run:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components config show
}}
Verify that the repository, ref, component name, command, schedule, and other settings are correct.
=== Check deployment history ===
Run:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components deployment list
}}
Use the deployment information to determine whether the build or deployment failed.
=== Check your CI pipeline ===
If pushing a change does not create a deployment, check your repository's CI pipeline.
Verify that:
* the CI job ran after the Git push;
* the CI job completed successfully;
* the credentials required to trigger a Toolforge deployment are configured correctly.
See [[Help:Toolforge/Deploy your tool#Triggering a deployment from your CI runner|Triggering a deployment from your CI runner]] for the current setup instructions.
=== Rebuild the source manually ===
For troubleshooting, you can ask the Components Service to rebuild the source even when a matching build already exists:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components deployment create --force-build
}}
This creates a manual deployment and bypasses the normal Push-to-Deploy workflow.
=== Restart an unchanged component manually ===
If you need to recreate a component even though its configuration has not changed:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components deployment create --force-run
}}
This also creates a manual deployment and is normally only needed for troubleshooting or administrative purposes.
== See also ==
* [[Help:Toolforge/Deploy your tool|Deploy your tool]]
* [[Help:Toolforge/Deploy your tool#Triggering a deployment from your CI runner|Triggering a deployment from your CI runner]]
* [[Help:Toolforge/Build Service|Toolforge Build Service]]
* [[Help:Toolforge/Jobs Framework|Toolforge Jobs Framework]]
bn3fdtrpzc6377tyt2qk753lrqpbcvp
2458710
2458709
2026-09-20T06:24:17Z
SSapaty (WMF)
28955
Fix ascii diagram
2458710
wikitext
text/x-wiki
= Migrating an existing Build Service job to Push-to-Deploy =
{{Note|Push-to-Deploy is currently a beta feature of Toolforge.}}
This tutorial explains how to migrate an existing Toolforge job that uses an image built with the [[Help:Toolforge/Build Service|Build Service]] to the Components Service, and then configure [[Help:Toolforge/Deploy your tool|Push-to-Deploy]] so that future Git pushes automatically trigger deployments.
If you already use the Build Service together with the Jobs Framework, most of the information needed to create the Components configuration can be generated automatically from your existing job.
This tutorial uses a scheduled job named <code>recent-changes</code> as an example.
The migration has two main stages:
# Migrate the existing Build Service job to the Components Service and verify that the component works.
# Configure your Git repository to trigger Components deployments automatically when you push changes.
The second stage enables Push-to-Deploy.
== Before you begin ==
This tutorial assumes that:
* you already have a working Toolforge job;
* the job uses an image built with the Toolforge Build Service;
* the source code is stored in a publicly accessible Git repository supported by the Build Service;
* you have access to the Toolforge tool account;
* you can commit and push changes to the source repository.
Become your tool before running Toolforge commands in this tutorial:
{{Codesample|lang=shell-session|scheme=light|code=
$ become my-tool
}}
Replace <code>my-tool</code> with the name of your Toolforge tool.
== Inspect your existing job ==
First, inspect the current Jobs Framework configuration:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge jobs list -o long
}}
For example, the <code>recent-changes</code> job looks similar to this:
{{Codesample|lang=text|scheme=light|code=
Job name: recent-changes
Command: recent-changes
Job type: scheduled: */5 * * * *
Image: tool-komla-test2/recent-changes:latest@sha256:...
File log: no
Emails: onfailure
Resources: mem: default, cpu: 0.5
Mounts: none
Retry: yes: 1 time(s)
Timeout: yes: 120s
}}
The important detail is the image:
{{Codesample|lang=text|scheme=light|code=
tool-komla-test2/recent-changes:latest
}}
An image whose name is under your tool's Build Service namespace, such as <code>tool-my-tool/my-image</code>, indicates that the job is using a Build Service image.
You can also inspect previous builds:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge build list
}}
For the example job, the relevant build points to:
{{Codesample|lang=text|scheme=light|code=
Source repository: https://gitlab.wikimedia.org/repos/cloud/wmcs/ptd-samples/tutorial-1
Image name: recent-changes
Git ref: main
}}
== Generate a Components configuration ==
The Components Service can generate an example configuration from your existing Jobs Framework jobs.
Run:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components config generate
}}
{{Warning|The generated configuration is an example. Review and validate it before deploying it.}}
For the <code>recent-changes</code> job, the command generates a configuration similar to:
{{Codesample|lang=yaml|scheme=light|name=toolforge.yaml|code=
config_version: v1beta1
components:
recent-changes:
build:
repository: https://gitlab.wikimedia.org/repos/cloud/wmcs/ptd-samples/tutorial-1
ref: main
run:
command: recent-changes
cpu: '0.5'
emails: onfailure
filelog: false
memory: 0.5Gi
mount: none
retry: 1
schedule: '*/5 * * * *'
timeout: 120
component_type: scheduled
}}
The generated configuration translates the existing Jobs Framework configuration into a Components Service configuration.
For example:
{| class="wikitable"
! Existing job setting
! Components configuration setting
|-
| Job name
| Component name
|-
| Build Service source repository
| <code>build.repository</code>
|-
| Git branch or ref
| <code>build.ref</code>
|-
| Job command
| <code>run.command</code>
|-
| CPU
| <code>run.cpu</code>
|-
| Memory
| <code>run.memory</code>
|-
| Email notification policy
| <code>run.emails</code>
|-
| File logging
| <code>run.filelog</code>
|-
| Filesystem mount
| <code>run.mount</code>
|-
| Retry count
| <code>run.retry</code>
|-
| Job timeout
| <code>run.timeout</code>
|-
| Cron schedule
| <code>run.schedule</code>
|-
| Job type
| <code>component_type</code>
|}
== Review the generated configuration ==
Do not deploy the generated configuration without reviewing it.
Check that at least the following values match your existing deployment:
* source repository;
* Git branch or ref;
* command;
* component type;
* CPU and memory;
* schedule, for scheduled jobs;
* retry and timeout settings;
* email notification settings;
* filesystem mount settings.
You should also verify that the generated component represents the job you intend to migrate.
For example:
{{Codesample|lang=yaml|scheme=light|code=
components:
recent-changes:
}}
corresponds to the existing Jobs Framework job named <code>recent-changes</code>.
If your tool contains several jobs, <code>toolforge components config generate</code> may generate configuration for multiple components. Review each generated component before continuing.
== Add the configuration to your source repository ==
The Components configuration can be hosted anywhere that is publicly accessible.
For this tutorial, save the generated configuration as <code>toolforge.yaml</code> in the same Git repository as the source code. Keeping the source code and deployment configuration together makes it easier to track changes to both in Git.
You can use any publicly accessible Git repository. For Toolforge projects, we recommend using a repository in the Wikimedia GitLab <code>toolforge-repos</code> namespace.
For example:
{{Codesample|lang=text|scheme=light|code=
my-project/
├── Procfile
├── requirements.txt
├── source-files...
└── toolforge.yaml
}}
Commit and push the file:
{{Codesample|lang=shell-session|scheme=light|code=
$ git add toolforge.yaml
$ git commit -m "Add Toolforge Components configuration"
$ git push
}}
The repository and ref in <code>toolforge.yaml</code> should point to the source you want Toolforge to build.
== Create the Components configuration ==
Once <code>toolforge.yaml</code> is publicly accessible, create the Components Service configuration from the copy stored in the repository.
For example:
{{Codesample|lang=shell-session|scheme=light|code=
$ curl 'https://gitlab.wikimedia.org/toolforge-repos/my-tool/-/raw/main/toolforge.yaml?ref_type=heads' {{!}} toolforge components config create
}}
Replace the URL with the public raw URL for your <code>toolforge.yaml</code> file.
This avoids having to clone the repository or download <code>toolforge.yaml</code> on the Toolforge bastion.
Inspect the stored configuration:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components config show
}}
Review the output and make sure it matches the configuration you committed.
== Test the migration with a manual deployment ==
Before configuring Push-to-Deploy, create a deployment manually to verify that the Components Service can successfully manage the workload.
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components deployment create
}}
{{Note|This is a manual Components Service deployment. Running <code>toolforge components deployment create</code> yourself is not Push-to-Deploy. Push-to-Deploy will be configured later in this tutorial so that your CI runner creates deployments automatically after you push changes to Git.}}
The Components Service will use the configuration to build and deploy the component.
You can optionally attach a description:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components deployment create --description "Test recent-changes migration"
}}
Normally, the Components Service can reuse an existing build when the source has not changed.
To force a new build, use:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components deployment create --force-build
}}
To force the component to run again even when its configuration has not changed, use:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components deployment create --force-run
}}
These options are normally unnecessary for the initial migration.
== Verify the migrated component ==
Inspect your Components deployments:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components deployment list
}}
You can also inspect the current Components configuration:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components config show
}}
For a scheduled component, verify that the resulting job exists:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge jobs list -o long
}}
Compare the resulting job with the configuration you reviewed earlier.
For the example, you would expect to see a job similar to:
{{Codesample|lang=text|scheme=light|code=
Job name: recent-changes
Command: recent-changes
Job type: scheduled: */5 * * * *
}}
Also verify that the job behaves as expected when its next scheduled execution occurs.
At this point, you have migrated the workload to the Components Service and verified that it works. The next step is to configure your Git repository so that future deployments are triggered automatically.
== Configure Push-to-Deploy ==
Push-to-Deploy allows a CI runner to trigger a Components Service deployment after changes are pushed to your Git repository.
Follow the instructions in [[Help:Toolforge/Deploy your tool#Triggering a deployment from your CI runner|Triggering a deployment from your CI runner]] to configure your repository.
The setup gives the CI runner the credentials it needs to request a Components deployment for your tool.
After the CI configuration has been added to the repository, future deployments can follow this workflow:
{{Codesample|lang=text|scheme=light|code=
Git push
↓
CI runner
↓
Components Service deployment
↓
Toolforge workload
}}
You no longer need to SSH to Toolforge and manually run <code>toolforge components deployment create</code> for normal deployments.
== Test Push-to-Deploy ==
After configuring your CI runner, make a small change to <code>toolforge.yaml</code>.
For example, change the schedule:
{{Codesample|lang=yaml|scheme=light|name=toolforge.yaml|code=
run:
schedule: '0 * * * *'
}}
Commit and push the change:
{{Codesample|lang=shell-session|scheme=light|code=
$ git add toolforge.yaml
$ git commit -m "Change recent-changes schedule"
$ git push
}}
Do not manually create a deployment.
The push should trigger your CI runner, which creates a Components Service deployment automatically.
Check your repository's CI pipeline and verify that it completes successfully.
You can then inspect the deployment from Toolforge:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components deployment list
}}
For the scheduled example, verify that the new schedule has been applied:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge jobs list -o long
}}
If the deployment was triggered automatically and the new configuration was applied, Push-to-Deploy is working.
== Avoid independently managing the migrated job ==
After migrating a scheduled or continuous job, do not maintain a second independently managed copy of the same workload.
The component created by the Components Service may appear when you inspect workloads using Jobs Framework commands. This does not mean that you should continue managing that workload separately with <code>toolforge jobs</code> commands.
After migration, make changes to the component through <code>toolforge.yaml</code> and your source repository.
== Make future changes through Git ==
Once Push-to-Deploy is configured, make changes to your source code or <code>toolforge.yaml</code>, commit them, and push them to the configured Git branch.
For example:
{{Codesample|lang=shell-session|scheme=light|code=
$ git add .
$ git commit -m "Update recent-changes"
$ git push
}}
The CI runner triggers the deployment. You do not normally need to manually rebuild the Build Service image, update the Jobs Framework job, or create a Components deployment from the Toolforge bastion.
This allows the Git repository to track the source code and, when you choose to keep them together, the Toolforge deployment configuration.
== What changed? ==
Before migration, the deployment workflow is approximately:
{{Codesample|lang=text|scheme=light|code=
Git repository
↓
Build Service
↓
Build Service image
↓
Jobs Framework job
}}
The maintainer manages the build and the job separately.
After migrating to the Components Service and configuring Push-to-Deploy, the workflow becomes:
{{Codesample|lang=text|scheme=light|code=
Git repository
├── source code
└── toolforge.yaml
↓
Git push
↓
CI runner
↓
Components Service deployment
├── build
└── run
↓
Toolforge workload
}}
The Components configuration describes how the workload should be built and run, while Push-to-Deploy allows changes pushed to Git to trigger deployment automatically.
== Troubleshooting ==
=== Check the generated configuration ===
If the generated configuration does not look correct, do not deploy it immediately.
Compare it with:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge jobs list -o long
$ toolforge build list
}}
Correct <code>toolforge.yaml</code> before creating the Components configuration.
=== Check the Components configuration ===
Run:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components config show
}}
Verify that the repository, ref, component name, command, schedule, and other settings are correct.
=== Check deployment history ===
Run:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components deployment list
}}
Use the deployment information to determine whether the build or deployment failed.
=== Check your CI pipeline ===
If pushing a change does not create a deployment, check your repository's CI pipeline.
Verify that:
* the CI job ran after the Git push;
* the CI job completed successfully;
* the credentials required to trigger a Toolforge deployment are configured correctly.
See [[Help:Toolforge/Deploy your tool#Triggering a deployment from your CI runner|Triggering a deployment from your CI runner]] for the current setup instructions.
=== Rebuild the source manually ===
For troubleshooting, you can ask the Components Service to rebuild the source even when a matching build already exists:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components deployment create --force-build
}}
This creates a manual deployment and bypasses the normal Push-to-Deploy workflow.
=== Restart an unchanged component manually ===
If you need to recreate a component even though its configuration has not changed:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components deployment create --force-run
}}
This also creates a manual deployment and is normally only needed for troubleshooting or administrative purposes.
== See also ==
* [[Help:Toolforge/Deploy your tool|Deploy your tool]]
* [[Help:Toolforge/Deploy your tool#Triggering a deployment from your CI runner|Triggering a deployment from your CI runner]]
* [[Help:Toolforge/Build Service|Toolforge Build Service]]
* [[Help:Toolforge/Jobs Framework|Toolforge Jobs Framework]]
f5thy2zlg0tfadnakxzy9uuat741f4t